Ultra-high-definition video stream multi-modal low-latency partitioning method based on dynamic resource scheduling
By deploying probe modules on edge computing nodes to build a terminal capability profile model and dynamically adjusting bitrate allocation and routing paths, low-latency and high-quality video stream transmission is achieved in a multi-terminal environment. This solves the problem of low-latency and high-quality video segmentation in the video stream transmission of multi-modal terminals in existing technologies, and achieves stable and smooth video playback in a multi-terminal environment.
Patent Information
- Application Number
- CN202511226560.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-29
- Publication Date
- 2025-11-18
- Estimated Expiration
- 2045-08-29
AI Technical Summary
Existing technologies have failed to effectively address the heterogeneous characteristics of multimodal terminals in ultra-high-definition video streaming, resulting in problems such as high transmission latency, inconsistent video quality, and playback stuttering. In particular, it is difficult to achieve low-latency and high-quality video segmentation in complex network environments.
By deploying probe modules on edge computing nodes, multimodal terminal data is collected in real time to build a terminal capability profile model. Combined with network status, the video stream is divided into high-motion and static background slices by dynamically adjusting bitrate allocation and routing paths, and the transmission path of video content is optimized to achieve low-latency and high-quality video stream transmission.
It enables low-latency, high-quality video streaming in multimodal terminal environments, improving the flexibility and accuracy of video streaming, reducing the probability of stuttering, and ensuring the stability and smoothness of video playback.
Smart Images

Figure CN120751136B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of video streaming, in particular to a multi-modal low-latency partitioning method for ultra-high-definition video streaming based on dynamic resource scheduling. BACKGROUND
[0002] With the rapid development of ultra-high-definition video technology, users have increasingly high requirements for video viewing experience, not only pursuing higher video clarity, but also having strict standards for real-time and smoothness of video transmission. Ultra-high-definition video streams have the characteristics of large data volume and high code rate, and when partitioning is performed on multi-modal terminals, there are problems such as large differences in terminal capabilities, complex and variable network environments, which can easily lead to high transmission delay, inconsistent video quality, and playback lag.
[0003] CN109685802B provides a low-latency video segmentation real-time preview method, mainly for single-device video segmentation preview scenarios, without considering the heterogeneous characteristics of multi-modal terminals. This method does not involve collecting and adapting terminal decoding computing power, rendering buffer saturation, and other capability parameters, and cannot adjust the processing strategy according to terminal differences. Its transmission mechanism does not distinguish between high motion-sensitive areas and static background areas of the video, and uses a single transmission path, which is difficult to meet the low-latency demand in complex network environments. At the same time, there is a lack of mechanisms for dynamically adjusting code rate allocation and routing paths based on network status, and there is no feasibility verification and negative feedback adjustment link on the terminal side, which cannot solve the problems of poor video quality consistency and large delay fluctuations in multi-terminal scenarios, and cannot meet the efficient partitioning needs of ultra-high-definition video streams on multi-modal terminals. SUMMARY
[0004] To solve the problems raised in the background art, the present application provides a multi-modal low-latency partitioning method for ultra-high-definition video streams based on dynamic resource scheduling, which can real-time perceive terminal capabilities and network status, and realize low-latency and high-quality partitioning of ultra-high-definition video streams on multi-modal terminals by dynamically adjusting code rate allocation and routing paths.
[0005] The technical solution of the present application is as follows:
[0006] According to an aspect of the present application, a multi-modal low-latency partitioning method for ultra-high-definition video streams based on dynamic resource scheduling is provided, comprising:
[0007] A probe module is deployed on an edge computing node to collect multi-modal terminal data in real time; based on the collected multi-modal terminal data, a terminal capability portrait model is constructed to determine the comprehensive capability index threshold of each terminal when maintaining a target frame rate in the future time domain, and a dynamically updated terminal capability portrait is generated;
[0008] By using the terminal capability profile, combining real-time network bandwidth and packet loss rate data, a double-target optimization model is established: the first target is to minimize the end-to-end transmission delay as the core, and the constraint condition is the terminal comprehensive capability index threshold; the second target is to maximize the consistency of multi-terminal video quality as the guide, and the constraint condition is the real-time transcoding resource margin of the edge node; the output is the optimal code rate allocation scheme of each terminal, as well as the low delay path and the high compression path.
[0009] The ultra-high definition video stream is divided into high motion sensitive area slices and static background slices, for high motion sensitive slices, the code rate is allocated according to the optimal code rate allocation scheme and transmitted by the low delay path, for static background slices, the code rate is allocated according to the optimal code rate allocation scheme and transmitted by the high compression path.
[0010] The decoding power simulation decoding time of the terminal capability profile is used to verify the slice stall risk, and the slice stall risk value is generated, based on which the double-target weight coefficient is adjusted, and the optimal code rate allocation scheme and the low delay path and the high compression path are regenerated.
[0011] As a further optimization of the present application, the probe module adopts a lightweight architecture and is embedded in the underlying system of the edge computing node, obtains the decoding power of the terminal device through the calling interface, monitors the filling state of the rendering buffer through the video rendering framework to evaluate the rendering capability, obtains the network jitter tolerance by measuring the data transmission delay fluctuation between the terminal and the edge node, and obtains the screen gamut parameters through the terminal display management interface.
[0012] As a further optimization of the present application, the multi-modal terminal feature extraction includes: calculating the decoder load rate, the rendering buffer saturation, the network jitter tolerance and the color gamut coverage of the terminal;
[0013] Wherein, the decoder load rate is the ratio of the actual running time of the decoder to the total running time, the rendering buffer saturation is the ratio of the current buffer size to the maximum capacity, the network jitter tolerance is the ratio of the network delay fluctuation to the average delay, and the color gamut coverage is the ratio of the terminal color gamut area to the standard color space color gamut area.
[0014] As a further optimization of the present application, the construction formula of the terminal capability profile model is:
[0015] ;
[0016] Wherein, is the comprehensive capability index, is the feature of decoding power, is the feature of rendering capability, is the feature of network jitter tolerance, Features of screen color gamut parameters For dynamic adjustment coefficients, The rate of performance change within the time window;
[0017] By weighting and fusing the static comprehensive capability index with the dynamically predicted comprehensive capability index threshold, a dynamically updated terminal capability profile is generated. The calculation formula is as follows:
[0018] ;
[0019] in, A profile index for terminal capabilities. The dynamic weighting coefficient is calculated as follows: , For adjustment coefficients, For terminal performance volatility, As a comprehensive capability index, The threshold for the comprehensive capability index;
[0020] The generated terminal capability profile is presented in the form of a multi-dimensional capability index, including: terminal capability profile index, characteristics of decoding computing power, characteristics of rendering capability, characteristics of network jitter tolerance, and characteristics of screen color gamut parameters.
[0021] As a further optimization of this application, the first objective function of the bi-objective optimization model is:
[0022] ;
[0023] in, Indicates end-to-end transmission delay. Indicates terminal The video stream transmission delay, Indicates the edge node to the terminal The latency introduced by transcoding the video stream, Indicates network link transmission delay. Total number of terminals;
[0024] The first objective function constraint is: the comprehensive capability index ≥ the terminal comprehensive capability index threshold;
[0025] The second objective function is:
[0026] ;
[0027] in, This indicates consistent video quality across multiple devices. The average bitrate across all terminals. Indicates terminal The difference between the bitrate and the average bitrate;
[0028] The constraints for the second objective function are: ,in, For the terminal The transcoding resources required for the video stream, This represents the total amount of available transcoding resources for edge nodes.
[0029] As a further optimization of this application, the step of generating the optimal bitrate allocation scheme includes:
[0030] The upper limit of the bitrate is determined based on the comprehensive capability index threshold of the terminal capability profile model;
[0031] A multi-objective optimization algorithm is used to iteratively generate candidate solutions, prioritizing those that meet the threshold constraint of the terminal's comprehensive capability index.
[0032] The weighting coefficients of the two objectives are dynamically adjusted. When the network latency fluctuates greatly, the transmission latency is reduced first. When the terminal decoding computing power fluctuates greatly, the consistency of video quality is guaranteed first.
[0033] As a further optimization of this application, the low-latency path selection for the high motion-sensitive region slice includes:
[0034] Prioritize 5G direct links or fiber optic transmission paths with the lowest end-to-end transmission latency.
[0035] Ensure sufficient path bandwidth with minimal fluctuations;
[0036] Forward error correction technology is introduced to reduce retransmission delays caused by network packet loss.
[0037] As a further optimization of this application, the high-compression path selection for the static background slice includes:
[0038] Prioritize transmission paths that support high compression encoding;
[0039] Utilize edge node caching technology to reduce bandwidth consumption from repeated transmissions;
[0040] Long-term reference frame technology is introduced to reuse static background frames to improve compression efficiency.
[0041] As a further optimization of this application, the formula for calculating the slice stuttering risk value is as follows:
[0042] ;
[0043] in, This indicates the risk value of slice stuttering. and These are the weighting coefficients for decoding time and rendering buffer saturation, respectively. Indicates the decoding time. This indicates the maximum allowed decoding time for the terminal. This indicates the saturation level of the render buffer. This represents the maximum allowed saturation level of the buffer.
[0044] As a further optimization of this application, the dynamic adjustment formula for the dual-objective weight coefficients is as follows:
[0045] ;
[0046] in, This represents the adjusted weighting coefficient. This represents the weighting coefficients before adjustment. To adjust the coefficient, This represents the risk value for slice stuttering.
[0047] The beneficial effects of this application are as follows:
[0048] Advantage (1): Multimodal data fusion and dynamic profiling improve the adaptability of resource allocation.
[0049] This application uses a probe module to collect real-time data on the terminal's decoding computing power, rendering capabilities, network jitter tolerance, and screen color gamut parameters, constructing a multi-dimensional terminal capability profile model to dynamically quantify the terminal's performance in different scenarios. This multi-modal data fusion mechanism can accurately perceive the terminal's real-time capability status and, combined with a time series prediction model, predict the terminal's comprehensive capability index threshold in the future time domain, thereby dynamically adjusting bitrate allocation and transmission paths. Compared to traditional static resource allocation strategies, this application can adapt to terminal performance fluctuations and network status changes, significantly improving the flexibility and accuracy of resource allocation and ensuring efficient matching between video stream transmission and terminal capabilities.
[0050] Advantage (2): The dual-objective optimization and content-aware segmentation strategy takes into account both transmission efficiency and image quality consistency.
[0051] By establishing a dual-objective optimization model centered on minimizing end-to-end latency and maximizing video quality consistency across multiple terminals, this application achieves a balance between multiple objectives under resource constraints. The model dynamically generates the optimal bitrate allocation scheme by combining terminal capability profiles and real-time network status, and divides the video stream into high-motion-sensitive regions and static background regions based on video content analysis. This content-aware differentiated transmission strategy not only reduces transmission latency in high-dynamic regions but also optimizes bandwidth utilization in static regions through high-compression coding, thereby balancing transmission efficiency and image quality consistency in a multi-terminal environment and improving the overall playback experience.
[0052] Advantage (3): The closed-loop feedback mechanism effectively copes with performance fluctuations and network state changes.
[0053] This application incorporates a feedback mechanism that dynamically assesses the stuttering risk of video stream segments through decoding time simulation and rendering buffer state verification. When a decoding timeout or buffer overflow risk exceeds a threshold, the system generates a negative feedback vector carrying resource gap parameters. This drives a bi-objective optimization model to dynamically adjust weight coefficients and bitrate allocation schemes, and triggers edge node routing path switching. This mechanism ensures that the system can respond in real time to sudden situations such as decreased terminal decoding capabilities and increased network jitter, proactively optimizing transmission strategies, significantly reducing the probability of stuttering, and improving the stability and smoothness of video playback. Attached Figure Description
[0054] Figure 1 This is a schematic diagram of the overall process of the multimodal low-latency partitioning method for ultra-high-definition video streams based on dynamic resource scheduling;
[0055] Figure 2 A detailed flowchart of the S100 steps of the ultra-high-definition video stream multimodal low-latency partitioning method based on dynamic resource scheduling;
[0056] Figure 3 A detailed flowchart of the S200 steps of the ultra-high-definition video stream multimodal low-latency partitioning method based on dynamic resource scheduling;
[0057] Figure 4 A detailed flowchart of the S300 step-by-step method for multimodal low-latency partitioning of ultra-high-definition video streams based on dynamic resource scheduling;
[0058] Figure 5 A detailed flowchart of the S400 step-by-step method for multimodal low-latency partitioning of ultra-high-definition video streams based on dynamic resource scheduling. Detailed Implementation
[0059] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.
[0060] In existing technologies, the partitioning of ultra-high-definition video streams typically employs fixed bitrate allocation and static routing path selection, making it difficult to dynamically adjust based on the real-time capabilities of the terminal and network conditions. For example, some terminals have limited decoding computing power and cannot handle high bitrate video streams, easily leading to stuttering due to decoding timeouts; furthermore, when network bandwidth fluctuates, fixed routing paths may fail to guarantee stable transmission latency, impacting user experience. Moreover, existing methods are insufficient in handling the consistency of video quality across multiple terminals, making it difficult to ensure a balanced level of video viewing quality across different terminals while meeting their latency requirements. To address these issues, please refer to [link to relevant documentation / reference]. Figure 1 This application illustrates a method for multimodal low-latency partitioning of ultra-high-definition video streams based on dynamic resource scheduling, according to an embodiment of this application. The method includes:
[0061] S100: Deploy probe modules on edge computing nodes to collect multimodal terminal data in real time; based on the collected multimodal terminal data, construct a terminal capability profile model, determine the comprehensive capability index threshold of each terminal when maintaining the target frame rate in the future time domain, and generate a dynamically updated terminal capability profile.
[0062] S200: Utilizing terminal capability profiles and combining real-time network bandwidth and packet loss rate data, a dual-objective optimization model is established: The first objective is to minimize end-to-end transmission latency, with the constraint being the threshold of the terminal's comprehensive capability index; the second objective is to maximize the consistency of video quality across multiple terminals, with the constraint being the real-time transcoding resource margin at edge nodes; the model outputs the optimal bitrate allocation scheme for each terminal, as well as low-latency paths and high-compression paths.
[0063] S300: Divides the ultra-high-definition video stream into high motion-sensitive area slices and static background slices. For high motion-sensitive slices, the bitrate is allocated according to the optimal bitrate allocation scheme and transmitted using a low-latency path. For static background slices, the bitrate is allocated according to the optimal bitrate allocation scheme and transmitted using a high-compression path.
[0064] Step S400: Simulate decoding time using the decoding computing power of the terminal capability profile, verify slice stuttering risk by combining the rendering buffer state, generate slice stuttering risk value, adjust the dual-target weight coefficients based on the slice stuttering risk value, regenerate the optimal bitrate allocation scheme, as well as the low-latency path and high-compression path.
[0065] In the multimodal low-latency partitioning method for ultra-high-definition video streams based on dynamic resource scheduling, S100 collects multimodal terminal data in real time by deploying probe modules on edge computing nodes. Combining terminal decoding computing power, rendering capabilities, network jitter tolerance, and screen color gamut characteristics, it predicts the maximum bitrate threshold that the terminal can carry to maintain the target frame rate in the future and constructs a dynamically updated terminal capability profile.
[0066] Please refer toFigure 2 The diagram illustrates a flowchart of an example of an ultra-high-definition video stream multimodal low-latency partitioning method S100 based on dynamic resource scheduling, the contents of which include:
[0067] S110: Deploy probe modules and collect multimodal terminal data.
[0068] The probe module is responsible for collecting multimodal terminal data in real time, including decoding computing power, rendering capabilities, network jitter tolerance, and screen color gamut parameters.
[0069] In one possible implementation, the probe module employs a lightweight architecture and is embedded in the underlying system of the edge computing node to minimize the consumption of system resources.
[0070] Specifically, the probe module monitors the terminal device's operating status and obtains decoding computing power through system call interfaces. It monitors the rendering buffer's fill status using the video rendering framework, recording buffer occupancy to assess the terminal device's rendering capabilities. The probe module measures data transmission latency fluctuations between the terminal device and edge computing nodes to obtain network jitter tolerance. Finally, the probe module utilizes the terminal device's display management interface to obtain screen color gamut parameters.
[0071] S120: Extract multimodal terminal features based on the collected multimodal terminal data.
[0072] Feature extraction is performed on the collected multimodal terminal performance data to further quantify the key performance indicators of the terminals, providing key data support for subsequent terminal capability profiling.
[0073] In one possible implementation, the decoder load rate of the terminal is calculated based on the characteristics of the decoding computing power, that is, the ratio of the actual running time of the decoder to the total running time, in order to evaluate its decoding capability.
[0074] In one possible implementation, the render buffer saturation is calculated based on the characteristics of the rendering capability, i.e., by the ratio of the current buffer size to the maximum capacity, to assess how close it is to saturation.
[0075] In one possible implementation, network jitter tolerance is calculated based on the characteristics of network jitter tolerance, that is, it is calculated by the ratio of the terminal's network latency fluctuation to the terminal's average network latency, in order to evaluate the network jitter tolerance characteristics.
[0076] In one possible implementation, the color gamut coverage of the terminal is calculated based on the characteristics of the screen color gamut parameters. This is done by calculating the ratio of the current color gamut area of the terminal to the color gamut area of the standard color space, in order to evaluate the characteristics of the screen color gamut parameters.
[0077] S130: Construct a terminal capability profile model.
[0078] Based on the multimodal terminal features extracted from S120, a terminal capability profile model is constructed to quantify the terminal's performance in different video playback scenarios. The terminal capability profile model integrates features of the terminal's decoding computing power, rendering buffer saturation, network jitter tolerance, and screen color gamut to construct a multidimensional data model that can dynamically reflect the terminal's capabilities.
[0079] In one possible implementation, the constructed terminal capability profile model is as follows:
[0080] ;
[0081] in, As a comprehensive capability index, Characteristics of decoding computing power, Features of rendering capabilities This refers to the characteristics of network jitter tolerance. Features of screen color gamut parameters For dynamic adjustment coefficients, This represents the rate of performance change within the time window.
[0082] S140: Determine the comprehensive capability index threshold of the terminal when maintaining the target frame rate in the future time domain through the terminal capability profile model.
[0083] Based on the terminal capability profile model constructed using S130, the comprehensive capability index threshold for maintaining the target frame rate in the future time domain is determined. Determining the comprehensive capability index threshold involves analyzing the terminal's decoding capabilities, rendering buffer state, network jitter tolerance, and screen color gamut characteristics, combined with the dynamic changing trends of multimodal terminal characteristics, to determine the comprehensive capability index threshold.
[0084] In one possible implementation, a time-series prediction model is used to predict the comprehensive capability index threshold in the future time domain, based on a terminal capability profile model. The model's input includes a sequence of historical comprehensive capability indices, multimodal terminal characteristics, and the target frame rate requirement. The output is the minimum comprehensive capability index threshold required to maintain the target frame rate in the future time domain. For example, if the target frame rate is 60fps, the model needs to predict whether the terminal can maintain a comprehensive capability index ≥ the minimum comprehensive capability index threshold within the next T seconds to ensure smooth video playback and stable image quality.
[0085] S150: Generate a profile of the terminal's capabilities.
[0086] Based on the terminal capability profile model constructed by S130 and the comprehensive capability index threshold predicted by S140, a multi-dimensional index system for quantifying the terminal capability status is constructed by integrating the static performance characteristics of the terminal and the dynamic prediction results, namely, the terminal capability profile.
[0087] In one possible implementation, the process for generating a terminal capability profile includes:
[0088] First, input the comprehensive capability index output by the terminal capability profile model of S130 and the comprehensive capability index threshold of S140.
[0089] By weighting and fusing the static comprehensive capability index with the dynamically predicted comprehensive capability index threshold, a dynamically updated terminal capability profile is generated. In one possible implementation, the following formula is used for calculation:
[0090] ;
[0091] in, A profile index for terminal capabilities. The dynamic weighting coefficient is calculated as follows: , For adjustment coefficients, For terminal performance volatility, As a comprehensive capability index, This is the threshold for the comprehensive capability index.
[0092] Output result:
[0093] The generated terminal capability profile is presented in the form of a multi-dimensional capability index, including: terminal capability profile index, characteristics of decoding computing power, characteristics of rendering capabilities, characteristics of network jitter tolerance, and characteristics of screen color gamut parameters.
[0094] In the multimodal low-latency partitioning method for ultra-high-definition video streams based on dynamic resource scheduling, S200 constructs a dual-objective optimization model by combining terminal capability profiles and real-time network status. The model aims to minimize end-to-end transmission latency and maximize the consistency of video quality across multiple terminals, and dynamically generates the optimal bitrate allocation scheme and the corresponding low-latency or high-compression routing path.
[0095] Please refer to Figure 3 The diagram illustrates a flowchart of an example of an ultra-high-definition video stream multimodal low-latency partitioning method S200 based on dynamic resource scheduling, the contents of which include:
[0096] S210: Establish a dual-objective optimization model.
[0097] Based on terminal capability profiles, real-time network bandwidth, and packet loss rate data, a dual-objective optimization model is established. This model comprises two core optimization objectives: the first objective is to minimize end-to-end transmission latency, and the second objective is to maximize the consistency of video quality across multiple terminals.
[0098] In one possible implementation, the dual-objective optimization model is constructed based on the following two core objectives:
[0099] The primary optimization objective is to minimize end-to-end transmission latency. Specifically, end-to-end transmission latency is composed of multiple factors, including the transmission latency of the video stream, the transcoding latency of edge nodes, and the transmission latency of the network link. This objective function can be expressed as:
[0100] ;
[0101] in, Indicates end-to-end transmission delay. Indicates terminal The video stream transmission delay, Indicates the edge node to the terminal The latency introduced by transcoding the video stream, Indicates network link transmission delay. This represents the total number of terminals.
[0102] The constraint on this objective is the minimum comprehensive capability index threshold for the terminal, namely: .
[0103] The second optimization objective is to maximize the consistency of video quality across multiple terminals. Specifically, video quality consistency refers to the similarity in image quality among the video streams received by all terminals. This objective function can be expressed as:
[0104] ;
[0105] in, This indicates consistent video quality across multiple devices. The average bitrate across all terminals. Indicates terminal The smaller the difference between the bitrate and the average bitrate, the higher the consistency of video quality.
[0106] The constraint on this objective is the real-time transcoding resource margin of the edge nodes, namely: ,in, For the terminal The transcoding resources required for the video stream, This represents the total amount of available transcoding resources for edge nodes.
[0107] Based on the first and second optimization objectives, in one possible implementation, the objective function of the bi-objective optimization model can be expressed as:
[0108] ;
[0109] in, Indicates end-to-end transmission delay. This indicates consistent video quality across multiple devices. It is a weighting coefficient used to adjust the relative importance between two objectives.
[0110] S220: Generates the optimal bitrate allocation scheme, as well as the corresponding low-latency path and high-compression path.
[0111] The optimal bitrate allocation scheme is based on the output of the dual-objective optimization model, combined with the terminal capability profile index, real-time network bandwidth and packet loss rate data, and solves the optimal bitrate allocation scheme for each terminal through a multi-objective optimization algorithm.
[0112] In one possible implementation, the steps for generating the optimal bitrate allocation scheme include:
[0113] Based on the comprehensive capability index threshold output by the terminal capability profile model, the upper limit of bitrate and the range of resource consumption for each terminal are determined.
[0114] A multi-objective optimization algorithm is used to iteratively calculate the initial population, generating multiple candidate bitrate allocation schemes. During the iteration process, priority is given to satisfying the minimum comprehensive capability index threshold constraint of the terminal, while balancing the goals of end-to-end transmission delay and video quality consistency.
[0115] The weighting coefficients of the dual-objective function are dynamically adjusted based on real-time network volatility and the stability of the terminal capability profile. For example, when network latency fluctuates significantly, priority is given to reducing transmission latency; when terminal decoding computing power fluctuates significantly, priority is given to ensuring video quality consistency.
[0116] The final bitrate allocation result is output. The generated optimal bitrate allocation scheme is presented in the form of a list, which includes the bitrate value of each terminal, the corresponding target frame rate, the resource consumption of edge nodes, and the path selection priority.
[0117] Based on the generated optimal bitrate allocation scheme, low-latency or high-compression paths are dynamically selected for each terminal, and the path status is monitored in real time to trigger dynamic switching.
[0118] In one possible implementation, the low-latency path employs a 5G direct link combined with forward error correction technology, suitable for scenarios with high decoding load or low network jitter tolerance. The high-compression path employs edge node caching and bitrate conversion technology, suitable for scenarios with low packet loss rate and sufficient decoding capability.
[0119] In the method for multimodal low-latency partitioning of ultra-high-definition video streams based on dynamic resource scheduling, S300 divides the ultra-high-definition video stream into high motion-sensitive regions and static background regions, dynamically allocates bitrates based on terminal capability profiles and real-time network status, and selects low-latency paths or high-compression paths to achieve low-latency and efficient partitioning of multimodal terminals.
[0120] Please refer to Figure 4 The flowchart illustrates an example of an ultra-high-definition video stream multimodal low-latency partitioning method S300 based on dynamic resource scheduling, the contents of which include:
[0121] S310: Divides the ultra-high-definition video stream into high motion-sensitive area slices and static background slices.
[0122] In the segmentation of ultra-high-definition video streams, the dynamic characteristics of the video content have a significant impact on transmission efficiency and playback experience. Therefore, video content analysis technology is used to divide the video stream into high motion-sensitive regions and static background regions to achieve differentiated processing for different regions. High motion-sensitive regions contain a large number of rapidly changing visual elements, while static background regions consist of relatively stable background information.
[0123] In one possible implementation, the division between high motion-sensitive region slices and static background slices relies on motion feature analysis of the video content. By identifying motion patterns between video frames, the video content is divided into high motion-sensitive regions and static background regions.
[0124] S320: Bitrate allocation and low-latency path selection in high motion-sensitive regions.
[0125] For high motion-sensitive areas, bitrate allocation is performed according to the optimal bitrate allocation scheme, and low-latency paths are selected for transmission. Since high motion-sensitive areas typically contain a large amount of dynamically changing information, their video data volume is large, placing high demands on transmission bandwidth and decoding performance. Therefore, by combining the terminal capability profile model and real-time network status, the bitrate allocation strategy is dynamically adjusted, and the optimal low-latency path is selected to ensure smooth playback of the video stream.
[0126] In one possible implementation, the bitrate allocation for high motion-sensitive regions is based on the decoding computing power, network jitter tolerance, and current network bandwidth in the terminal capability profile model, ensuring that the terminal can stably decode and render video streams under low-latency paths.
[0127] Specifically, the steps for determining high motion-sensitive regions according to the optimal bitrate allocation scheme include:
[0128] Analyze the terminal's overall capability index and predict whether the terminal can maintain an overall capability index greater than or equal to the minimum overall capability index threshold within the next T seconds;
[0129] Based on the decoding computing power characteristics of the terminal, determine its maximum bit rate.
[0130] The bitrate allocation is dynamically adjusted based on the saturation level of the terminal's rendering buffer. If the terminal's rendering buffer is close to saturation, the system can reduce the bitrate to reduce rendering latency; if the buffer load is low, the bitrate can be appropriately increased to optimize image quality.
[0131] For terminals with low tolerance for network jitter, lower bitrate video streams are prioritized to reduce the risk of transmission instability caused by network fluctuations.
[0132] High-motion-sensitive regions are extremely sensitive to transmission latency; therefore, low-latency paths are selected to ensure the real-time performance of the video stream. The selection of low-latency paths is based on the following criteria:
[0133] Prioritize the path with the lowest end-to-end transmission latency, such as 5G direct links or fiber optic transmission, to reduce the data transmission time in the network.
[0134] Ensure sufficient bandwidth and minimal fluctuations in the path to avoid video stuttering due to insufficient bandwidth.
[0135] Introducing forward error correction technology into low-latency paths reduces retransmission delays caused by network packet loss by adding redundant data, thereby improving transmission stability.
[0136] S330: Bitrate allocation and high compression path selection for static background areas.
[0137] For static background areas, the bitrate is allocated according to the optimal bitrate allocation scheme, and a high-compression path is selected for transmission.
[0138] Since static background areas typically contain less dynamic information, their video data volume is relatively small, resulting in lower requirements for transmission bandwidth and decoding performance. Therefore, by combining terminal capability profiling models and real-time network conditions, the bitrate allocation strategy is dynamically adjusted, and the optimal high-compression path is selected to improve resource utilization and transmission efficiency.
[0139] In one possible implementation, the bitrate allocation for the static background area is based on the rendering capabilities, screen color gamut parameters, and current network bandwidth in the terminal capability profile model.
[0140] Specifically, the steps for determining the static background region according to the optimal bitrate allocation scheme include:
[0141] Analyze the terminal's overall capability index and predict whether the terminal can maintain an overall capability index greater than or equal to the minimum overall capability index threshold within the next T seconds;
[0142] The bitrate allocation is dynamically adjusted based on the saturation of the terminal's rendering buffer. If the terminal's rendering buffer is close to saturation, the bitrate is reduced to decrease rendering latency; if the buffer load is low, the bitrate can be appropriately increased to optimize image quality.
[0143] For terminals with a high tolerance for network jitter, lower bitrate video streams are prioritized to reduce the risk of transmission instability caused by network fluctuations.
[0144] Static background slicing requires high bandwidth efficiency; therefore, a high-compression path is selected to optimize transmission efficiency. The selection of the high-compression path is based on the following criteria:
[0145] Prioritize transmission paths that support high compression encoding to reduce data transmission volume.
[0146] Ensure the path supports edge node caching technology to reduce bandwidth consumption caused by repeated transmissions.
[0147] Introducing long-term reference frame technology into the high-compression path reduces data transmission volume by reusing static background frames, thereby improving compression efficiency.
[0148] In the multimodal low-latency partitioning method for ultra-high-definition video streams based on dynamic resource scheduling, S400 quantifies the risk value of segmentation stuttering by simulating decoding time consumption and verifying rendering buffer status based on terminal capability profiles, dynamically adjusts the weight coefficients of the dual-objective optimization model, and regenerates the optimal bitrate allocation scheme and low-latency / high-compression path to achieve real-time optimization of resource scheduling and transmission strategies to ensure the smoothness of video streams and consistency of image quality.
[0149] Please refer to Figure 5 The flowchart illustrates an example of an ultra-high-definition video stream multimodal low-latency partitioning method S400 based on dynamic resource scheduling, the contents of which include:
[0150] S410: Simulation of decoding time based on terminal capability profile.
[0151] The decoding capability of a terminal is a key factor affecting the smoothness of video playback and the stability of image quality. Therefore, based on the terminal capability profile, the decoding time of the video stream is simulated to evaluate whether the terminal can meet the target frame rate requirement under current resource conditions.
[0152] In one possible implementation, the simulation formula for decoding time consumption is as follows:
[0153] ;
[0154] in, Indicates the decoding time. For the bitrate of the video stream, The data size of a single video frame. This refers to the decoding computing power of the terminal. This formula determines the decoding time by quantifying the relationship between the terminal's decoding capability and the amount of data in the video stream.
[0155] S420: Verify the risk of slice stuttering based on the decoding time and rendering buffer status of the terminal capability profile, and generate slice stuttering risk values.
[0156] Based on the decoding time simulation, the S420 further incorporates the terminal's rendering buffer status to assess the risk of stuttering in video stream segments. The terminal device's rendering buffer stores the video frames to be played, and its fill status directly affects the smoothness of video playback. The S420 monitors the terminal's rendering buffer status in real time to obtain the rendering buffer's saturation level.
[0157] In one possible implementation, the formula for assessing slice stuttering risk is:
[0158] ;
[0159] in, This indicates the risk value of slice stuttering. and These are the weighting coefficients for decoding time and rendering buffer saturation, respectively. Indicates the decoding time. This indicates the maximum allowed decoding time for the terminal. This indicates the saturation level of the render buffer. This represents the maximum allowed saturation level of the buffer.
[0160] If the terminal's decoding computing power is low, causing the decoding time to exceed the maximum allowable value, then A larger value indicates a higher risk of stuttering. Similarly, if the render buffer saturation is high, close to the maximum allowed value, then... The value will also be large, indicating a risk of overflow.
[0161] After completing the slice lag risk assessment, the system adjusts its resource scheduling strategy based on the slice lag risk value.
[0162] S430: Based on the slice stuttering risk value, the dual-objective weight coefficients are dynamically adjusted to regenerate the optimal bitrate allocation scheme, as well as low-latency paths and high-compression paths.
[0163] The weight coefficients of the objective function in the bi-objective optimization model determine the balance between end-to-end transmission latency and multi-terminal video quality consistency. When the slice stuttering risk value indicates a high risk of stuttering on the terminal, the weight coefficients of the bi-objective model need to be dynamically adjusted to optimize resource scheduling strategies and ensure smooth playback and stable image quality of the video stream.
[0164] In one possible implementation, the dynamic adjustment of the bi-objective weighting coefficients is calculated based on the following formula:
[0165] ;
[0166] in, This represents the adjusted weighting coefficient. This represents the weighting coefficients before adjustment. To adjust the coefficient, This represents the risk value for slice stuttering.
[0167] After determining the new weight coefficients, the objective function of the bi-objective optimization model is updated, and the optimal bitrate allocation scheme, as well as the corresponding low-latency path and high-compression path, are regenerated using the S220 step.
[0168] The S100-S400 system, based on dynamic resource scheduling, achieves low-latency segmentation and transmission optimization of ultra-high-definition video streams through multi-modal terminal data acquisition and terminal capability profiling. Its core innovation lies in: utilizing probe modules to collect real-time data on the terminal's decoding computing power, rendering capabilities, network jitter tolerance, and screen color gamut parameters to construct a dynamic terminal capability profiling model and predict the bitrate threshold that the terminal can handle. Combined with a dual-objective optimization model, it dynamically generates the optimal bitrate allocation scheme and low-latency / high-compression transmission paths; through video content analysis, it divides the video stream into high-motion-sensitive areas and static background areas, and introduces a feedback mechanism to verify stuttering risks based on decoding time simulation and rendering buffer status verification, dynamically adjusting and optimizing model weights and resource scheduling strategies. The advantages of this application are: (1) It achieves accurate perception of terminal performance through multimodal data fusion and dynamic profiling, thereby improving the adaptability of resource allocation; (2) The dual-objective optimization and content-aware partitioning strategy takes into account both transmission efficiency and image quality consistency, thereby reducing end-to-end latency; (3) The closed-loop feedback mechanism effectively responds to terminal performance fluctuations and network status changes, thereby significantly reducing the probability of stuttering; and finally, it achieves high-quality, low-latency ultra-high-definition video transmission in a multi-terminal, multi-network environment.
[0169] The basic principles of this application have been described above with reference to specific embodiments. However, it should be noted that the advantages, benefits, and effects mentioned in this application are merely examples and not limitations, and should not be considered as essential features of each embodiment of this application. Furthermore, the specific details disclosed above are for illustrative and facilitative purposes only, and are not limitations. These details do not limit the application to the necessity of employing the aforementioned specific details for implementation.
[0170] The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use this application. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein can be applied to other aspects without departing from the scope of this application. Therefore, this application is not intended to be limited to the aspects shown herein, but rather to be carried out within the widest scope consistent with the principles and novel features of the invention herein.
[0171] The above description is merely a preferred embodiment of this application and is not intended to limit this application. Any modifications or equivalent substitutions made within the spirit and principles of this application should be included within the protection scope of this application.
Claims
1. A method for multimodal low-latency partitioning of ultra-high-definition video streams based on dynamic resource scheduling, characterized in that, include: Deploy probe modules on edge computing nodes to collect multimodal terminal data in real time; Based on the collected multimodal terminal data, a terminal capability profile model is constructed to determine the comprehensive capability index threshold of each terminal when maintaining the target frame rate in the future time domain, and to generate a dynamically updated terminal capability profile. By leveraging terminal capability profiles and combining real-time network bandwidth and packet loss rate data, a dual-objective optimization model is established: the first objective is to minimize end-to-end transmission latency, with the constraint being the threshold of the terminal's comprehensive capability index; the second objective is to maximize the consistency of video quality across multiple terminals, with the constraint being the real-time transcoding resource margin at edge nodes; the model outputs the optimal bitrate allocation scheme for each terminal, as well as the low-latency path and the high-compression path. The ultra-high-definition video stream is divided into high motion-sensitive area slices and static background slices. For high motion-sensitive slices, the bitrate is allocated according to the optimal bitrate allocation scheme and transmitted using a low-latency path. For static background slices, the bitrate is allocated according to the optimal bitrate allocation scheme and transmitted using a high-compression path. The decoding time is simulated by using the decoding computing power of the terminal capability profile. The slice stuttering risk is verified by combining the rendering buffer state. The slice stuttering risk value is generated. Based on the slice stuttering risk value, the dual-objective weight coefficient is adjusted, and the optimal bitrate allocation scheme, as well as the low-latency path and high-compression path, are regenerated.
2. The method for multimodal low-latency partitioning of ultra-high-definition video streams based on dynamic resource scheduling according to claim 1, characterized in that, The probe module adopts a lightweight architecture and is embedded in the underlying system of the edge computing node. It obtains the decoding computing power of the terminal device by calling the interface, monitors the filling status of the rendering buffer through the video rendering framework to evaluate the rendering capability, obtains the network jitter tolerance by measuring the data transmission latency fluctuation between the terminal and the edge node, and obtains the screen color gamut parameters through the terminal display management interface.
3. The method for multimodal low-latency partitioning of ultra-high-definition video streams based on dynamic resource scheduling according to claim 1, characterized in that, The multimodal terminal data extraction includes: the decoder load rate of the computing terminal, the rendering buffer saturation, the network jitter tolerance, and the color gamut coverage. Among them, decoder load rate is the ratio of actual decoder running time to total running time, rendering buffer saturation is the ratio of current buffer size to maximum capacity, network jitter tolerance is the ratio of network latency fluctuation to average latency, and color gamut coverage is the ratio of terminal color gamut area to standard color space color gamut area.
4. The method for multimodal low-latency partitioning of ultra-high-definition video streams based on dynamic resource scheduling according to claim 1, characterized in that, The formula for constructing the terminal capability profile model is as follows: ; in, As a comprehensive capability index, Characteristics of decoding computing power, Features of rendering capabilities This refers to the characteristics of network jitter tolerance. The color gamut parameter is a characteristic of the screen, and α is a dynamic adjustment coefficient. The rate of performance change within the time window; By weighting and fusing the static comprehensive capability index with the dynamically predicted comprehensive capability index threshold, a dynamically updated terminal capability profile is generated. The calculation formula is as follows: ; in, The terminal capability profiling index, where β is a dynamic weighting coefficient, is calculated as follows: , For adjustment coefficients, For terminal performance volatility, As a comprehensive capability index, The threshold for the comprehensive capability index; The generated terminal capability profile is presented in the form of a multi-dimensional capability index, including: terminal capability profile index, characteristics of decoding computing power, characteristics of rendering capability, characteristics of network jitter tolerance, and characteristics of screen color gamut parameters.
5. The method for multimodal low-latency partitioning of ultra-high-definition video streams based on dynamic resource scheduling according to claim 1, characterized in that, The first objective function of the bi-objective optimization model is: ; in, Indicates end-to-end transmission delay. Indicates terminal The video stream transmission delay, Indicates the edge node to the terminal The latency introduced by transcoding the video stream, Indicates network link transmission delay. Total number of terminals; The first objective function constraint is: the comprehensive capability index ≥ the terminal comprehensive capability index threshold; The second objective function is: ; in, This indicates consistent video quality across multiple devices. The average bitrate across all terminals. Indicates terminal The difference between the bitrate and the average bitrate; The constraints for the second objective function are: ,in, For the terminal The transcoding resources required for the video stream, This represents the total amount of available transcoding resources for edge nodes.
6. The method for multimodal low-latency partitioning of ultra-high-definition video streams based on dynamic resource scheduling according to claim 1, characterized in that, The steps for generating the optimal bitrate allocation scheme include: The upper limit of the bitrate is determined based on the comprehensive capability index threshold of the terminal capability profile model; A multi-objective optimization algorithm is used to iteratively generate candidate solutions, prioritizing those that meet the threshold constraint of the terminal's comprehensive capability index. The weighting coefficients of the two objectives are dynamically adjusted. When the network latency fluctuates greatly, the transmission latency is reduced first. When the terminal decoding computing power fluctuates greatly, the consistency of video quality is guaranteed first.
7. The method for multimodal low-latency partitioning of ultra-high-definition video streams based on dynamic resource scheduling according to claim 1, characterized in that, The low-latency path selection for the highly motion-sensitive region slice includes: Prioritize 5G direct links or fiber optic transmission paths with the lowest end-to-end transmission latency. Ensure sufficient path bandwidth with minimal fluctuations; Forward error correction technology is introduced to reduce retransmission delays caused by network packet loss.
8. The method for multimodal low-latency partitioning of ultra-high-definition video streams based on dynamic resource scheduling according to claim 1, characterized in that, The high-compression path selection for the static background slice includes: Prioritize transmission paths that support high compression encoding; Utilize edge node caching technology to reduce bandwidth consumption from repeated transmissions; Long-term reference frame technology is introduced to reuse static background frames to improve compression efficiency.
9. The method for multimodal low-latency partitioning of ultra-high-definition video streams based on dynamic resource scheduling according to claim 1, characterized in that, The formula for calculating the slice stuttering risk value is as follows: ; in, This indicates the risk value of slice stuttering. and These are the weighting coefficients for decoding time and rendering buffer saturation, respectively. Indicates the decoding time. This indicates the maximum allowed decoding time for the terminal. This indicates the saturation of the render buffer. This represents the maximum allowed saturation level of the buffer.
10. The method for multimodal low-latency partitioning of ultra-high-definition video streams based on dynamic resource scheduling according to claim 9, characterized in that, The dynamic adjustment formula for the dual-objective weight coefficients is as follows: ; in, This represents the adjusted weighting coefficient. This represents the weighting coefficients before adjustment. To adjust the coefficient, This represents the risk value for slice stuttering.
Citation Information
Patent Citations
A low-latency video segmentation real-time preview method
CN109685802B
Video code rate adaptive method based on reinforcement learning for edge cellular network
CN116016987A
Inference acceleration method for real-time video stream application in computing network integration environment
CN116339977A