Global live broadcast bandwidth saving method based on region-of-interest optimization and related device

By performing region-of-interest analysis and stabilization on real-time video streams, combined with regional hierarchical coding and multi-cloud cost scheduling, the problems of high bandwidth cost and unstable coding quality in existing technologies are solved, achieving efficient bandwidth utilization and improved video service reliability.

CN121334409APending Publication Date: 2026-01-13HAIJIAO CLOUD (SHENZHEN) INFORMATION TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511523698.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-23
Publication Date
2026-01-13

AI Technical Summary

Technical Problem

In existing technologies, global high bitrate coding strategies lead to a linear increase in cross-border bandwidth costs. Traditional region of interest (ROI) coding technologies lack synergistic optimization with edge computing architectures, adaptive bitrate technologies, and global network scheduling. Simple ROI detection algorithms are prone to frequent changes in region identification results in dynamic scenarios, resulting in unstable coding quality and a lack of effective fault tolerance and degradation mechanisms, which affect the reliability and economic benefits of video services.

Method used

By performing region of interest (ROI) analysis on real-time video streams, identifying and stabilizing them, a stable ROI is obtained. This ROI is then used for hierarchical coding of the ROI, generating adaptive bitstreams with multiple quality levels. Transmission paths are dynamically allocated based on network bandwidth. By combining edge-side ROI processing, ABR bitrate adaptation, and multi-cloud cost scheduling, an intelligent closed-loop system is constructed to achieve precise allocation of coding resources and fault-tolerant rollback.

Benefits of technology

It improves bandwidth utilization efficiency, ensures a smooth transition in quality across focused areas, enhances viewing comfort and visual experience continuity, optimizes overall service quality and cost-effectiveness, and strengthens system robustness and availability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121334409A_ABST
    Figure CN121334409A_ABST
Patent Text Reader

Abstract

The embodiment of the invention relates to the technical field of video processing, and discloses a global live broadcast bandwidth saving method and device based on region-of-interest optimization, computer equipment and a computer readable storage medium. Identifying a region of interest in each frame of video image in the real-time video stream; performing stabilization processing on the region of interest in each frame of video image to obtain the region of interest after the stabilization processing; performing regional hierarchical coding on the real-time video stream based on the stabilized region of interest to obtain hierarchical coded video data; encapsulating the hierarchical coding video data into a self-adaptive code stream containing a plurality of quality gears; and distributing the adaptive code stream to a corresponding transmission path according to the quality gear. By means of the mode, the continuity and stability of visual experience are effectively improved, and the bandwidth utilization efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of video processing technology, specifically to a method, apparatus, computer device, and computer-readable storage medium for saving global live streaming bandwidth based on region of interest optimization. Background Technology

[0002] With the rapid growth of global ultra-high-definition video live streaming services, the requirements for video quality are also constantly increasing.

[0003] However, existing technologies for video processing have the following problems: First, to ensure image quality, a global high bitrate encoding strategy is generally adopted, which leads to a linear increase in cross-border bandwidth costs with the scale of users, resulting in poor economic benefits.

[0004] Secondly, traditional region of interest coding techniques often exist as isolated functions, lacking end-to-end collaborative optimization with edge computing architectures, adaptive bitrate techniques, and global network scheduling, making it difficult to achieve the expected results in real-world complex network environments.

[0005] More importantly, simple ROI detection algorithms are prone to frequent changes in region identification results in dynamic scenes, leading to unstable encoding quality and severely impairing the viewing experience. Furthermore, existing systems generally lack effective fault tolerance and degradation mechanisms when edge computing resources are limited or malfunctions, resulting in insufficient service reliability. Summary of the Invention

[0006] In view of the above problems, embodiments of the present invention provide a method, apparatus, computer device and computer-readable storage medium for saving global live streaming bandwidth based on region of interest optimization, to solve the problem of poor video quality caused by encoding and bandwidth in the prior art.

[0007] According to one aspect of the present invention, a global live streaming bandwidth saving method based on region of interest optimization is provided, the method comprising: Perform region of interest analysis on the received real-time video stream to identify the region of interest in each frame of the real-time video stream; The region of interest in each frame of video image is stabilized to obtain a stabilized region of interest; the stabilization process is used to spatially expand and temporally smooth the boundary of the region of interest. Based on the stabilized region of interest, the real-time video stream is subjected to regional hierarchical coding to obtain hierarchically coded video data; wherein the image quality parameters of the stabilized region of interest are better than those of the unstable region of interest. The graded coded video data is encapsulated into an adaptive bitstream containing multiple quality levels; where the higher the quality level of the adaptive bitstream, the larger the area of ​​the region of interest based on the stabilized processing of the video image or the higher the image quality parameter of the video image. Based on the quality level, the adaptive bitstream is allocated to the corresponding transmission path.

[0008] In one alternative approach, the step of performing region of interest analysis on the received real-time video stream to identify the region of interest in each frame of the real-time video stream further includes: Multimodal region of interest detection is performed on the video image to obtain visual detection results; Receive the marking information of the focal region coordinates in the video image from the broadcast control system; Receive user interaction behavior statistics; Based on the visual detection results, the labeling information, and the interaction behavior statistics, the region of interest in the video image and the detection confidence of the region of interest in the video image are determined.

[0009] In one optional approach, the visual detection result includes a detection confidence level; the stabilization process for the region of interest in each frame of video image to obtain a stabilized region of interest further includes: Determine whether the detection confidence of N consecutive video frames is greater than an entry threshold; When the threshold is exceeded, the region of interest in consecutive frames of the image is determined. After the region of interest takes effect, the detection confidence of the video image in subsequent frames continues to be detected; Determine whether the detection confidence of the subsequent M consecutive video images is less than the exit threshold; When the region of interest (ROI) takes effect, if the effective duration is longer than a preset time and the detection confidence of the subsequent M consecutive video frames is less than the exit threshold, the ROI in the subsequent M consecutive video frames is determined to be invalid.

[0010] In one alternative approach, when the threshold is exceeded, determining the region of interest in consecutive frames of images takes effect, including: Morphological dilation is performed on the region of interest in the images of consecutive video frames to cover the minute movement of the target in the region of interest, thereby expanding the space. A first-order low-pass filter or Kalman filter is applied to the center coordinates and size of the bounding box of the region of interest between consecutive video frames to smooth the motion trajectory and avoid box jumps, thereby achieving temporal smoothing.

[0011] In one optional approach, the stabilized region of interest includes a primary region of interest and a secondary region of interest; based on the stabilized region of interest, the real-time video stream is subjected to region-level coding to obtain coded video data; wherein the image quality parameters of the stabilized region of interest are superior to those of the unstable region of interest, including: For the region of main interest, set it as the first image quality parameter; For the secondary region of interest, set it as the second image quality parameter; For the background region, set it to the third image quality parameter; Based on the first image quality parameter, the second image quality parameter, and the third image quality parameter; Among them, the values ​​of the first image quality parameter, the second image quality parameter, and the third image quality parameter increase sequentially.

[0012] In one alternative approach, allocating the adaptive bitstream to the corresponding transmission path based on the quality level includes: When the current network bandwidth is within the first threshold range, the quality level of the adaptive bitstream is determined to be high, and the region of interest of the video image after stabilization is adopted in extended mode, and the first image quality parameter is adopted based on the region of interest after stabilization. When the current network bandwidth is in the second threshold range, the quality level of the adaptive bitstream is determined to be high, and the standard mode is adopted for the region of interest of the video image after stabilization processing, and the second image quality parameter is adopted for the region of interest after stabilization processing. When the current network bandwidth is in the third threshold range, the quality level of the adaptive bitstream is determined to be high, and the core mode is adopted for the region of interest of the video image after stabilization processing, and the third image quality parameter is adopted for the region of interest after stabilization processing.

[0013] According to another aspect of the present invention, a global live streaming bandwidth saving device based on region of interest optimization is provided, comprising: The recognition module is used to perform region of interest analysis on the received real-time video stream and identify the region of interest in each frame of the real-time video stream. The stabilization module is used to stabilize the region of interest in each frame of video image to obtain a stabilized region of interest; the stabilization process is used to spatially expand and temporally smooth the boundary of the region of interest. The encoding module is used to perform regional hierarchical encoding on the real-time video stream based on the stabilized region of interest to obtain hierarchically encoded video data; wherein the image quality parameters of the stabilized region of interest are better than those of the unstable region of interest. The encapsulation module is used to encapsulate hierarchically coded video data into an adaptive bitstream containing multiple quality levels; wherein, in an adaptive bitstream with a higher quality level, the area of ​​the region of interest based on the stabilized processing of the video image is larger or the image quality parameter of the video image is larger. The allocation module is used to allocate the adaptive bitstream to the corresponding transmission path according to the quality level.

[0014] In one alternative approach, the step of performing region of interest analysis on the received real-time video stream to identify the region of interest in each frame of the real-time video stream further includes: Multimodal region of interest detection is performed on the video image to obtain visual detection results; Receive the marking information of the focal region coordinates in the video image from the broadcast control system; Receive user interaction behavior statistics; Based on the visual detection results, the labeling information, and the interaction behavior statistics, the region of interest in the video image and the detection confidence of the region of interest in the video image are determined.

[0015] According to another aspect of the present invention, a computer device is provided, including: a processor, a memory, a communication interface, and a communication bus, wherein the processor, the memory, and the communication interface communicate with each other through the communication bus; The memory is used to store at least one executable instruction that causes the processor to perform the operation of the global live streaming bandwidth saving method based on region of interest optimization.

[0016] According to another aspect of the present invention, a computer-readable storage medium is provided, the storage medium storing at least one executable instruction, which, when executed on a computer device, causes the computer device to perform the operation of the global live streaming bandwidth saving method based on region of interest optimization.

[0017] This invention, through region-of-interest (ROI) analysis of the received real-time video stream, identifies the ROI in each frame of the video stream. The ROI in each frame is then stabilized to obtain a stabilized ROI. Based on the stabilized ROI, the real-time video stream undergoes region-level coding to obtain tiered coded video data. This tiered coded video data is then encapsulated into an adaptive bitstream containing multiple quality levels. The adaptive bitstream is then allocated to the corresponding transmission path according to the quality levels. By dynamically reallocating coding resources from visually non-critical regions to critical regions, "precise bitrate allocation" is achieved, resulting in a substantial improvement in bandwidth utilization efficiency.

[0018] By innovating the ROI stabilization mechanism, the flickering problem caused by simple ROI detection is fundamentally solved, ensuring a smooth transition of the quality of the focus area, greatly improving viewing comfort and immersion, and enhancing the continuity and stability of the visual experience.

[0019] Furthermore, this invention deeply couples edge-side ROI processing, ABR bitrate adaptation, and multi-cloud cost scheduling to construct an intelligent closed-loop system that perceives content, network, and economy. This not only optimizes individual technical indicators but also achieves the optimal balance between global service quality and cost-effectiveness from video source to user terminal. The fault-tolerant rollback and phased recovery mechanism based on image group boundaries ensures that the system has the ability to self-repair and gracefully degrade when facing fluctuations in computing resources or module anomalies, guaranteeing the extremely high availability required for commercial live streaming services and significantly enhancing system robustness and service availability.

[0020] The above description is merely an overview of the technical solutions of the embodiments of the present invention. In order to better understand the technical means of the embodiments of the present invention and to implement them in accordance with the contents of the specification, and to make the above and other objects, features and advantages of the embodiments of the present invention more apparent and understandable, specific embodiments of the present invention are described below. Attached Figure Description

[0021] The accompanying drawings are for illustrative purposes only and are not intended to limit the invention. Furthermore, the same reference numerals denote the same parts throughout the drawings. In the drawings: Figure 1 A flowchart illustrating the global live streaming bandwidth saving method based on region of interest optimization provided in an embodiment of the present invention is shown. Figure 2 This illustration shows a schematic diagram of the region of interest (ROI) in the global live streaming bandwidth saving method based on ROI optimization provided in an embodiment of the present invention. Figure 3This diagram illustrates the stabilization of the global live streaming bandwidth saving method based on region of interest optimization provided in an embodiment of the present invention. Figure 4 A flowchart illustrating a global live streaming bandwidth saving method based on region of interest optimization according to another embodiment of the present invention is shown. Figure 5 This diagram illustrates the structure of a global live streaming bandwidth saving device based on region of interest optimization provided in an embodiment of the present invention. Figure 6 A schematic diagram of the structure of a computer device provided in an embodiment of the present invention is shown. Detailed Implementation

[0022] Exemplary embodiments of the invention will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the invention are shown in the drawings, it should be understood that the invention can be implemented in various forms and should not be limited to the embodiments set forth herein.

[0023] Figure 1 A flowchart illustrating a global live streaming bandwidth saving method based on region of interest optimization according to an embodiment of the present invention is shown. This method is executed by a computer device. The computer device can be an edge node of a distributed network. Figure 1 As shown, the method includes the following steps: Step 110: Perform region of interest analysis on the received real-time video stream to identify the region of interest in each frame of the real-time video stream.

[0024] The Region of Interest (ROI) is the area that the user is interested in.

[0025] In this embodiment of the invention, ROI analysis integrates data from at least two of the following information sources: computer vision detection results of the video frames themselves, marking information from the broadcasting system, and interactive behavior statistics from user terminals.

[0026] Specifically, multimodal region of interest (ROI) detection is performed on the video image to obtain visual detection results. This involves receiving the video stream from the source and routing it to the nearest edge processing node in the network topology. A lightweight neural network is used. Specifically, MobileNetV3-SSD or YOLOv5n models can be employed to detect faces, bodies, and common objects (such as balls and goods) in real time. When computational power is limited, a Haar cascade classifier combined with optical flow can be used for motion region detection. Simultaneously, the system receives the marking information of the focal region coordinates in the video image from the production system and fuses it with the visual detection results. The focal region coordinates in the video image can be specific product close-ups, for example, when the production system's marked region overlaps with the visually detected face region, the confidence level of that region is increased to the highest. User interaction behavior statistics, such as attention and comments on certain areas in historical video images, are received to determine regions that the user may be interested in. Based on the visual detection results, the marking information, and the interaction behavior statistics, the ROI in the video image and the detection confidence level of the ROI in the video image are determined. In this embodiment of the invention, the region of interest may include a primary region of interest and a secondary region of interest, each with a corresponding detection confidence level.

[0027] Step 120: Stabilize the region of interest in each frame of video image to obtain the stabilized region of interest.

[0028] The stabilization process is used to spatially expand and temporally smooth the boundaries of the region of interest. For example... Figure 2 As shown in this embodiment of the invention, during the stabilization process, an entry threshold, an exit threshold, and a minimum dwell time (i.e., a preset duration) are set. Specifically, it is determined whether the detection confidence of N consecutive video frames is greater than the entry threshold; when it is greater than the entry threshold, the region of interest (ROI) in the consecutive frames is determined to be active; after the ROI is active, the detection confidence of subsequent frames is detected; it is determined whether the detection confidence of the subsequent M consecutive video frames is less than the exit threshold; when the ROI is active, the active duration is greater than the preset duration and the detection confidence of the subsequent M consecutive video frames is less than the exit threshold, the ROI in the subsequent M consecutive video frames is determined to be inactive. Morphological dilation is performed on the ROI in the consecutive video frames to cover the minute movements of the target within the ROI, thereby expanding the spatial area; a first-order low-pass filter or Kalman filter is applied to the center coordinates and size of the bounding box of the ROI between consecutive video frames to smooth the motion trajectory and avoid box jumps, thereby smoothing the temporal area. In this embodiment of the invention, the stabilized region of interest includes a primary region of interest and a secondary region of interest.

[0029] In one embodiment of the present invention, the threshold hysteresis is set as follows: the "entry threshold" of the ROI is set to 0.8, and the "exit threshold" is set to 0.4. The ROI is considered to be officially effective only when the detection confidence score of 3 consecutive frames is higher than 0.8; after it is effective, the ROI is considered to have disappeared only when the confidence score of 5 consecutive frames is lower than 0.4. This mechanism effectively filters out instantaneous false detections and missed detections.

[0030] Once the ROI takes effect, the system initializes a tracker (such as KCF, SORT, or DeepSORT) to follow the target. This tracker outputs either "association confidence" or "motion matching score".

[0031] Minimum dwell time: Once an ROI takes effect, its state must remain unchanged for at least 1 second (30 frames at 30fps). During this period, even if the confidence level of a single frame or a few frames falls below the exit threshold, the ROI state will not change to prevent flickering caused by brief occlusion or changes in lighting.

[0032] Spatial smoothing: Performs morphological dilation on the boundaries of a defined ROI region, expanding outward by 5%-10% of pixels to cover minute movements of the target.

[0033] Temporal smoothing: A first-order low-pass filter or Kalman filter is applied to the center coordinates and size of the ROI bounding box between consecutive frames to smooth the motion trajectory and avoid box jumps. The smoothed confidence score is calculated as: α × Current frame detection confidence score + (1-α) × Previous frame smoothed confidence score. Here, α is a smoothing factor between 0 and 1 (e.g., 0.3). The smaller α is, the greater the weight of historical data, the smoother the result, and the stronger the resistance to instantaneous fluctuations.

[0034] Step 130: Based on the stabilized region of interest, perform region-level coding on the real-time video stream to obtain coded video data.

[0035] The image quality parameters of the region of interest after stabilization are better than those of the region of interest after destabilization.

[0036] For the region of main interest, set it as the first image quality parameter; For the secondary region of interest, set it as the second image quality parameter; For the background region, set it to the third image quality parameter; Based on the first image quality parameter, the second image quality parameter, and the third image quality parameter; The values ​​of the first, second, and third image quality parameters increase sequentially. A higher image quality parameter indicates lower image quality.

[0037] Specifically, taking the widely supported H.265 / HEVC encoder as an example, ROI mask information is passed in through its provided region encoding or ROI Map interface. The parameter settings are as follows: For the primary ROI region (region of interest), the quantization parameter (QP) is set to 22, enabling advanced quantization tools such as RDOQ and using more complex motion estimation search algorithms. For secondary ROI regions: the quantization parameter is set to 28, using standard motion search. For background regions: the quantization parameter is set to 38 or higher, which can disable some time-consuming encoding tools and even downsample the chroma components. To avoid a surge in overall computational overhead caused by ROI encoding, a fast encoding preset is used for the background region, while a slower preset is used for the ROI region, keeping the overall encoding time within an acceptable range, thus balancing complexity and quality. Simultaneously, to ensure keyframe consistency, the ROI region strategy is forcibly recalculated and applied at IDR frames or scene transition frames, ensuring that the visual quality of the new scene starts from a consistent starting point and avoiding quality jumps across GOPs caused by ROI switching.

[0038] Step 140: Encapsulate the graded coded video data into an adaptive bitstream containing multiple quality levels.

[0039] In adaptive bitstreams with higher quality levels, the area of ​​the region of interest based on the stabilized processing of the video image is larger, or the image quality parameters of the video image are higher.

[0040] In this embodiment of the invention, a dynamic correlation mechanism between bitrate level and ROI intensity is implemented during adaptive bitstream generation. This is the scope of this application. Specifically, multi-quality level generation involves using an encoder (e.g., x265) or transcoder (e.g., FFmpeg) to generate multiple ABR levels for the same input source. For example, three levels are generated: 1080p@3Mbps, 720p@1.5Mbps, and 540p@800kbps.

[0041] Step 150: Allocate the adaptive bitstream to the corresponding transmission path according to the quality level.

[0042] When the current network bandwidth is within the first threshold range, the quality level of the adaptive bitstream is determined to be high, and the region of interest of the video image after stabilization is adopted in extended mode, and the first image quality parameter is adopted based on the region of interest after stabilization. When the current network bandwidth is in the second threshold range, the quality level of the adaptive bitstream is determined to be high, and the standard mode is adopted for the region of interest of the video image after stabilization processing, and the second image quality parameter is adopted for the region of interest after stabilization processing. When the current network bandwidth is in the third threshold range, the quality level of the adaptive bitstream is determined to be high, and the core mode is adopted for the region of interest of the video image after stabilization processing, and the third image quality parameter is adopted for the region of interest after stabilization processing.

[0043] By setting multiple quality levels to match different network bandwidths, bandwidth consumption can be effectively reduced, resource utilization can be improved, and video can be made smoother.

[0044] In one embodiment of the present invention, the ROI intensity adaptive strategy is as follows: For the high ROI (3Mbps): At this ROI level, the ROI area uses "extended mode". For example, the ROI area covers the entire person and their surrounding environment. Optimal encoding parameters (e.g., QP=20) are used for this area.

[0045] For the mid-range (1.5Mbps): Shrink the ROI area to "Standard Mode," for example, focusing on faces and upper bodies. Adjust quality parameters appropriately (e.g., QP=25).

[0046] For lower bitrates (800kbps): the ROI area is further shrunk to a "core mode," for example, retaining only key facial features. At the same time, the quality parameters of secondary ROIs (such as subtitles) can be reduced to be similar to the background, concentrating the bitrate to ensure the most essential information (e.g., QP=30).

[0047] Smooth transition guarantee: When switching gears at the boundaries of a segment or GOP, the system sets up a brief transition window. Within this window, the area or quality parameters of the ROI region are smoothly interpolated to avoid visual abrupt changes in the image due to gear switching.

[0048] Terminal adaptive logic: The standard HLS / DASH player automatically selects the appropriate level based on the current network bandwidth. The advantage of this embodiment is that users with good network conditions enjoy a wide range of high-quality data, while users with poor network conditions do not see a globally blurred mosaic, but rather a clear core focus combined with a highly compressed background. This results in a superior subjective experience compared to traditional ABR strategies even under poor network conditions.

[0049] In this embodiment of the invention, a multi-cloud scheduling module is also used to select paths based on cost and performance.

[0050] Specifically, the scheduling system continuously collects real-time metric data from edge nodes and paths globally, including performance and cost metrics. Performance metrics include end-to-end latency, jitter, and packet loss rate. Cost metrics can be calculated based on contracts with cloud service providers and operators, determining the equivalent transmission cost per unit of traffic at the current moment.

[0051] A weighted scoring algorithm is used to calculate a comprehensive score for each candidate path: Overall score = (1 - Delayed Normalized Value) + (1 - packet loss rate) + (1 - normalized cost); in, γ is a configurable weighting coefficient that reflects different business preferences for performance and cost. For example, cost-sensitive businesses can increase γ.

[0052] To prevent network oscillations caused by frequent path switching, a scheduling stabilization mechanism is introduced. Once a path is selected, it must be used for at least 30 seconds. A switching rate limit is also implemented, such as a maximum of two path switching attempts per minute. Switching is only triggered when the overall score of a candidate path is consistently higher than that of the current path, the difference exceeds a certain threshold, and the aforementioned time constraints are met.

[0053] In this embodiment of the invention, a fault-tolerant rollback and recovery mechanism is also provided.

[0054] When an anomaly detection metric is detected, the fault-tolerant control module immediately sends a "rollback" command to the encoding module. Specifically, the anomaly detection metrics include the following: Performance anomaly: The processing time of a single frame in the ROI analysis module continues to exceed the 150ms threshold for 10 consecutive frames.

[0055] Quality anomaly: The highest confidence score of the ROI detection model output is consistently (e.g., for 15 consecutive frames) below 0.2.

[0056] Resource anomaly: The CPU utilization of the edge node is consistently above 85% (e.g., for 30 seconds).

[0057] When any of the above indicators are triggered, the fault-tolerant control module immediately sends a "rollback" command to the encoding module.

[0058] After the current GOP (Group of Video Frames) encoding is completed, the system immediately switches to a preset safe mode encoding scheme. This scheme uses uniform and moderate encoding parameters (e.g., global QP=26) to ensure basic image quality. Switching at GOP boundaries ensures the continuity and decodeability of the video stream, avoiding decoder stuttering or screen tearing caused by instantaneous parameter changes.

[0059] Continue monitoring of abnormal metrics. Once all metrics return to normal (e.g., CPU utilization drops below 70% and model confidence returns to normal) and remain stable for more than a preset "recovery observation period" (e.g., 10 seconds), trigger the recovery process. The recovery process is gradual: the system first re-enables ROI analysis at a lower frequency (e.g., processing 1 frame every 5 frames) and adopts a conservative ROI strategy (e.g., narrowing the detection range). Subsequently, the frequency is gradually increased and full functionality is restored, adhering to the GOP boundary switching principle throughout. During recovery, the minimum dwell time and threshold hysteresis mechanism remain in effect to prevent the system from oscillating at the stability edge.

[0060] This invention, through region-of-interest (ROI) analysis of the received real-time video stream, identifies the ROI in each frame of the video stream. The ROI in each frame is then stabilized to obtain a stabilized ROI. Based on the stabilized ROI, the real-time video stream undergoes region-level coding to obtain tiered coded video data. This tiered coded video data is then encapsulated into an adaptive bitstream containing multiple quality levels. The adaptive bitstream is then allocated to the corresponding transmission path according to the quality levels. By dynamically reallocating coding resources from visually non-critical regions to critical regions, "precise bitrate allocation" is achieved, resulting in a substantial improvement in bandwidth utilization efficiency.

[0061] By innovating the ROI stabilization mechanism, the flickering problem caused by simple ROI detection is fundamentally solved, ensuring a smooth transition of the quality of the focus area, greatly improving viewing comfort and immersion, and enhancing the continuity and stability of the visual experience.

[0062] Furthermore, this invention deeply couples edge-side ROI processing, ABR bitrate adaptation, and multi-cloud cost scheduling to construct an intelligent closed-loop system that perceives content, network, and economy. This not only optimizes individual technical indicators but also achieves the optimal balance between global service quality and cost-effectiveness from video source to user terminal. The fault-tolerant rollback and phased recovery mechanism based on image group boundaries ensures that the system has the ability to self-repair and gracefully degrade when facing fluctuations in computing resources or module anomalies, guaranteeing the extremely high availability required for commercial live streaming services and significantly enhancing system robustness and service availability.

[0063] like Figure 4 As shown, in another embodiment of the present invention, the process specifically includes: Video stream access and edge routing: Receive video streams from the source and route them to the nearest edge processing node in the network topology.

[0064] Stabilization extraction of ROI regions: At the edge node, multimodal analysis is performed on the video frame to identify candidate regions of interest; stabilization processing is applied to the identification results, wherein the stabilization processing includes: applying entry threshold and exit threshold to form a threshold hysteresis effect, applying minimum dwell time constraint, and spatially expanding and temporally smoothing filtering the ROI region boundary, thereby outputting a temporally continuous and spatially stable ROI mask sequence.

[0065] Region-based coding based on ROI: Based on the stable ROI mask, different coding parameter sets are configured for different regions within the video frame, wherein image quality parameters are assigned to ROI regions that are superior to those to non-ROI regions.

[0066] Generate an adaptive bitstream associated with ROI intensity: Encapsulate the graded encoded video data into an adaptive bitstream containing multiple quality levels; where higher bitrate levels are associated with larger ROI areas or better ROI quality parameters, while lower bitrate levels are associated with smaller ROI areas or moderately reduced ROI quality parameters.

[0067] Cost-aware multi-cloud path scheduling: Based on real-time collected network performance metrics and link cost data, the optimal distribution path is selected for the generated adaptive bitstream.

[0068] Quality monitoring and feedback optimization: Collect terminal playback quality data and dynamically adjust ROI analysis strategies or bitrate configurations accordingly.

[0069] Fault tolerance handling in abnormal states: When a system abnormality is detected, it automatically reverts to the global unified coding mode; after the system recovers and stabilizes, the ROI hierarchical coding function is gradually re-enabled.

[0070] Global live streaming bandwidth savings based on region of interest optimization Figure 5 A schematic diagram of the structure of the global live streaming bandwidth saving device based on region of interest optimization provided in an embodiment of the present invention is shown. Figure 5 As shown, the device 500 includes: The recognition module 510 is used to perform region of interest analysis on the received real-time video stream and identify the region of interest in each frame of the real-time video stream. The stabilization module 520 is used to perform stabilization processing on the region of interest in each frame of video image to obtain a stabilized region of interest; the stabilization processing is used to perform spatial expansion and temporal smoothing filtering on the boundary of the region of interest. The encoding module 530 is used to perform regional hierarchical encoding on the real-time video stream based on the stabilized region of interest to obtain hierarchically encoded video data; wherein the image quality parameters of the stabilized region of interest are better than the image quality parameters of the unstable region of interest. The encapsulation module 540 is used to encapsulate hierarchically coded video data into an adaptive bitstream containing multiple quality levels; wherein, in an adaptive bitstream with a higher quality level, the area of ​​the region of interest based on the stabilized processing of the video image is larger or the image quality parameter of the video image is larger. The allocation module 550 is used to allocate the adaptive bitstream to the corresponding transmission path according to the quality level.

[0071] The specific working process of the global live streaming bandwidth saving device based on region of interest optimization in this embodiment of the invention is largely the same as the specific implementation steps of the method in the foregoing embodiments, and will not be repeated here.

[0072] This invention, through region-of-interest (ROI) analysis of the received real-time video stream, identifies the ROI in each frame of the video stream. The ROI in each frame is then stabilized to obtain a stabilized ROI. Based on the stabilized ROI, the real-time video stream undergoes region-level coding to obtain tiered coded video data. This tiered coded video data is then encapsulated into an adaptive bitstream containing multiple quality levels. The adaptive bitstream is then allocated to the corresponding transmission path according to the quality levels. By dynamically reallocating coding resources from visually non-critical regions to critical regions, "precise bitrate allocation" is achieved, resulting in a substantial improvement in bandwidth utilization efficiency.

[0073] By innovating the ROI stabilization mechanism, the flickering problem caused by simple ROI detection is fundamentally solved, ensuring a smooth transition of the quality of the focus area, greatly improving viewing comfort and immersion, and enhancing the continuity and stability of the visual experience.

[0074] Furthermore, this invention deeply couples edge-side ROI processing, ABR bitrate adaptation, and multi-cloud cost scheduling to construct an intelligent closed-loop system that perceives content, network, and economy. This not only optimizes individual technical indicators but also achieves the optimal balance between global service quality and cost-effectiveness from video source to user terminal. The fault-tolerant rollback and phased recovery mechanism based on image group boundaries ensures that the system has the ability to self-repair and gracefully degrade when facing fluctuations in computing resources or module anomalies, guaranteeing the extremely high availability required for commercial live streaming services and significantly enhancing system robustness and service availability.

[0075] Figure 6The diagram shows a structural schematic of a computer device provided in an embodiment of the present invention. The specific embodiments of the present invention do not limit the specific implementation of the computer device.

[0076] like Figure 6 As shown, the computer device may include: a processor 402, a communications interface 404, a memory 406, and a communications bus 408.

[0077] The processor 402, communication interface 404, and memory 406 communicate with each other via communication bus 408. Communication interface 404 is used to communicate with other network elements such as clients or other servers. The processor 402 executes program 410, specifically performing the relevant steps described above in the computer method embodiment.

[0078] Specifically, program 410 may include program code, which includes computer-executable instructions.

[0079] Processor 402 may be a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement embodiments of the present invention. The computer device includes one or more processors, which may be processors of the same type, such as one or more CPUs; or processors of different types, such as one or more CPUs and one or more ASICs.

[0080] Memory 406 is used to store program 410. Memory 406 may include high-speed RAM memory, and may also include non-volatile memory, such as at least one disk storage device.

[0081] Specifically, program 410 can be called by processor 402 to cause the computer device to perform the following operations: Perform region of interest analysis on the received real-time video stream to identify the region of interest in each frame of the real-time video stream; The region of interest in each frame of video image is stabilized to obtain a stabilized region of interest; the stabilization process is used to spatially expand and temporally smooth the boundary of the region of interest. Based on the stabilized region of interest, the real-time video stream is subjected to regional hierarchical coding to obtain hierarchically coded video data; wherein the image quality parameters of the stabilized region of interest are better than those of the unstable region of interest. The graded coded video data is encapsulated into an adaptive bitstream containing multiple quality levels; where the higher the quality level of the adaptive bitstream, the larger the area of ​​the region of interest based on the stabilized processing of the video image or the higher the image quality parameter of the video image. Based on the quality level, the adaptive bitstream is allocated to the corresponding transmission path.

[0082] In one alternative approach, the step of performing region of interest analysis on the received real-time video stream to identify the region of interest in each frame of the real-time video stream further includes: Multimodal region of interest detection is performed on the video image to obtain visual detection results; Receive the marking information of the focal region coordinates in the video image from the broadcast control system; Receive user interaction behavior statistics; Based on the visual detection results, the labeling information, and the interaction behavior statistics, the region of interest in the video image and the detection confidence of the region of interest in the video image are determined.

[0083] In one optional approach, the visual detection result includes a detection confidence level; the stabilization process for the region of interest in each frame of video image to obtain a stabilized region of interest further includes: Determine whether the detection confidence of N consecutive video frames is greater than an entry threshold; When the threshold is exceeded, the region of interest in consecutive frames of the image is determined. After the region of interest takes effect, the detection confidence of the video image in subsequent frames continues to be detected; Determine whether the detection confidence of the subsequent M consecutive video images is less than the exit threshold; When the region of interest (ROI) takes effect, if the effective duration is longer than a preset time and the detection confidence of the subsequent M consecutive video frames is less than the exit threshold, the ROI in the subsequent M consecutive video frames is determined to be invalid.

[0084] In one alternative approach, when the threshold is exceeded, determining the region of interest in consecutive frames of images takes effect, including: Morphological dilation is performed on the region of interest in the images of consecutive video frames to cover the minute movement of the target in the region of interest, thereby expanding the space. A first-order low-pass filter or Kalman filter is applied to the center coordinates and size of the bounding box of the region of interest between consecutive video frames to smooth the motion trajectory and avoid box jumps, thereby achieving temporal smoothing.

[0085] In one optional approach, the stabilized region of interest includes a primary region of interest and a secondary region of interest; based on the stabilized region of interest, the real-time video stream is subjected to region-level coding to obtain coded video data; wherein the image quality parameters of the stabilized region of interest are superior to those of the unstable region of interest, including: For the region of main interest, set it as the first image quality parameter; For the secondary region of interest, set it as the second image quality parameter; For the background region, set it to the third image quality parameter; Based on the first image quality parameter, the second image quality parameter, and the third image quality parameter; Among them, the values ​​of the first image quality parameter, the second image quality parameter, and the third image quality parameter increase sequentially.

[0086] In one alternative approach, allocating the adaptive bitstream to the corresponding transmission path based on the quality level includes: When the current network bandwidth is within the first threshold range, the quality level of the adaptive bitstream is determined to be high, and the region of interest of the video image after stabilization is adopted in extended mode, and the first image quality parameter is adopted based on the region of interest after stabilization. When the current network bandwidth is in the second threshold range, the quality level of the adaptive bitstream is determined to be high, and the standard mode is adopted for the region of interest of the video image after stabilization processing, and the second image quality parameter is adopted for the region of interest after stabilization processing. When the current network bandwidth is in the third threshold range, the quality level of the adaptive bitstream is determined to be high, and the core mode is adopted for the region of interest of the video image after stabilization processing, and the third image quality parameter is adopted for the region of interest after stabilization processing.

[0087] This invention provides a computer-readable storage medium storing at least one executable instruction that, when executed on a computer device, causes the computer device to perform the global live streaming bandwidth saving method based on region of interest optimization in any of the above method embodiments.

[0088] Executable instructions can be used to cause computer devices to perform the following operations: Perform region of interest analysis on the received real-time video stream to identify the region of interest in each frame of the real-time video stream; The region of interest in each frame of video image is stabilized to obtain a stabilized region of interest; the stabilization process is used to spatially expand and temporally smooth the boundary of the region of interest. Based on the stabilized region of interest, the real-time video stream is subjected to regional hierarchical coding to obtain hierarchically coded video data; wherein the image quality parameters of the stabilized region of interest are better than those of the unstable region of interest. The graded coded video data is encapsulated into an adaptive bitstream containing multiple quality levels; where the higher the quality level of the adaptive bitstream, the larger the area of ​​the region of interest based on the stabilized processing of the video image or the higher the image quality parameter of the video image. Based on the quality level, the adaptive bitstream is allocated to the corresponding transmission path.

[0089] In one alternative approach, the step of performing region of interest analysis on the received real-time video stream to identify the region of interest in each frame of the real-time video stream further includes: Multimodal region of interest detection is performed on the video image to obtain visual detection results; Receive the marking information of the focal region coordinates in the video image from the broadcast control system; Receive user interaction behavior statistics; Based on the visual detection results, the labeling information, and the interaction behavior statistics, the region of interest in the video image and the detection confidence of the region of interest in the video image are determined.

[0090] In one optional approach, the visual detection result includes a detection confidence level; the stabilization process for the region of interest in each frame of video image to obtain a stabilized region of interest further includes: Determine whether the detection confidence of N consecutive video frames is greater than an entry threshold; When the threshold is exceeded, the region of interest in consecutive frames of the image is determined. After the region of interest takes effect, the detection confidence of the video image in subsequent frames continues to be detected; Determine whether the detection confidence of the subsequent M consecutive video images is less than the exit threshold; When the region of interest (ROI) takes effect, if the effective duration is longer than a preset time and the detection confidence of the subsequent M consecutive video frames is less than the exit threshold, the ROI in the subsequent M consecutive video frames is determined to be invalid.

[0091] In one alternative approach, when the threshold is exceeded, determining the region of interest in consecutive frames of images takes effect, including: Morphological dilation is performed on the region of interest in the images of consecutive video frames to cover the minute movement of the target in the region of interest, thereby expanding the space. A first-order low-pass filter or Kalman filter is applied to the center coordinates and size of the bounding box of the region of interest between consecutive video frames to smooth the motion trajectory and avoid box jumps, thereby achieving temporal smoothing.

[0092] In one optional approach, the stabilized region of interest includes a primary region of interest and a secondary region of interest; based on the stabilized region of interest, the real-time video stream is subjected to region-level coding to obtain coded video data; wherein the image quality parameters of the stabilized region of interest are superior to those of the unstable region of interest, including: For the region of main interest, set it as the first image quality parameter; For the secondary region of interest, set it as the second image quality parameter; For the background region, set it to the third image quality parameter; Based on the first image quality parameter, the second image quality parameter, and the third image quality parameter; Among them, the values ​​of the first image quality parameter, the second image quality parameter, and the third image quality parameter increase sequentially.

[0093] In one alternative approach, allocating the adaptive bitstream to the corresponding transmission path based on the quality level includes: When the current network bandwidth is within the first threshold range, the quality level of the adaptive bitstream is determined to be high, and the region of interest of the video image after stabilization is adopted in extended mode, and the first image quality parameter is adopted based on the region of interest after stabilization. When the current network bandwidth is in the second threshold range, the quality level of the adaptive bitstream is determined to be high, and the standard mode is adopted for the region of interest of the video image after stabilization processing, and the second image quality parameter is adopted for the region of interest after stabilization processing. When the current network bandwidth is in the third threshold range, the quality level of the adaptive bitstream is determined to be high, and the core mode is adopted for the region of interest of the video image after stabilization processing, and the third image quality parameter is adopted for the region of interest after stabilization processing.

[0094] The algorithms or displays provided herein are not inherently related to any particular computer, virtual system, or other device. Various general-purpose systems can also be used in conjunction with the teachings herein. The required structure for constructing such systems is apparent from the above description. Furthermore, the embodiments of the present invention are not directed to any particular programming language. It should be understood that the content of the invention described herein can be implemented using various programming languages, and the above description of specific languages ​​is for the purpose of disclosing the best mode of implementation of the invention.

[0095] Numerous specific details are set forth in the specification provided herein. However, it will be understood that embodiments of the invention may be practiced without these specific details. In some instances, well-known methods, structures, and techniques have not been shown in detail so as not to obscure the understanding of this specification.

[0096] Similarly, it should be understood that, in order to streamline the invention and aid in understanding one or more of the various aspects of the invention, features of the embodiments of the invention are sometimes grouped together in a single embodiment, figure, or description thereof in the above description of exemplary embodiments of the invention. However, this disclosure should not be construed as reflecting an intention that the claimed invention requires more features than are expressly recited in each claim.

[0097] Those skilled in the art will understand that modules in the device of the embodiments can be adaptively changed and placed in one or more devices different from that embodiment. Modules, units, or components in the embodiments can be combined into a single module, unit, or component, and can be divided into multiple sub-modules, sub-units, or sub-components. Except where at least some of such features and / or processes or units are mutually exclusive, any combination can be used to combine all features disclosed in this specification (including the accompanying claims, abstract, and drawings) and all processes or units of any method or device so disclosed. Unless expressly stated otherwise, each feature disclosed in this specification (including the accompanying claims, abstract, and drawings) may be replaced by an alternative feature that serves the same, equivalent, or similar purpose.

[0098] It should be noted that the above embodiments are illustrative of the invention and not restrictive, and that those skilled in the art can devise alternative embodiments without departing from the scope of the appended claims. In the claims, any reference signs placed between parentheses should not be construed as limiting the claims. The word "comprising" does not exclude the presence of elements or steps not listed in the claims. The word "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. The invention can be implemented by means of hardware comprising several different elements and by means of a suitably programmed computer. In the unit claims enumerating several means, several of these means may be embodied by the same item of hardware. The use of the words first, second, and third, etc., does not indicate any order. These words can be interpreted as names. The steps in the above embodiments, unless otherwise specified, should not be construed as limiting the order of execution.

Claims

1. A global live streaming bandwidth saving method based on region of interest optimization, characterized in that, The method includes: Perform region of interest analysis on the received real-time video stream to identify the region of interest in each frame of the real-time video stream; The region of interest in each frame of video image is stabilized to obtain a stabilized region of interest; the stabilization process is used to spatially expand and temporally smooth the boundary of the region of interest. Based on the stabilized region of interest, the real-time video stream is subjected to regional hierarchical coding to obtain hierarchically coded video data; wherein the image quality parameters of the stabilized region of interest are better than those of the unstable region of interest. The graded coded video data is encapsulated into an adaptive bitstream containing multiple quality levels; where the higher the quality level of the adaptive bitstream, the larger the area of ​​the region of interest based on the stabilized processing of the video image or the higher the image quality parameter of the video image. Based on the quality level, the adaptive bitstream is allocated to the corresponding transmission path.

2. The method according to claim 1, characterized in that, The step of performing region of interest (ROI) analysis on the received real-time video stream to identify the ROI in each frame of the video stream further includes: Multimodal region of interest detection is performed on the video image to obtain visual detection results; Receive the marking information of the focal region coordinates in the video image from the broadcast control system; Receive user interaction behavior statistics; Based on the visual detection results, the labeling information, and the interaction behavior statistics, the region of interest in the video image and the detection confidence of the region of interest in the video image are determined.

3. The method according to claim 2, characterized in that, The step of stabilizing the region of interest in each frame of video image to obtain a stabilized region of interest further includes: Determine whether the detection confidence of N consecutive video frames is greater than an entry threshold; When the threshold is exceeded, the region of interest in consecutive frames of the image is determined. After the region of interest is activated, the detection confidence of the video image in subsequent frames continues to be detected; Determine whether the detection confidence of the subsequent M consecutive video images is less than the exit threshold; When the region of interest (ROI) takes effect, if the effective duration is longer than a preset time and the detection confidence of the subsequent M consecutive video frames is less than the exit threshold, the ROI in the subsequent M consecutive video frames is determined to be invalid.

4. The method according to any one of claims 1-3, characterized in that, When the threshold is exceeded, the region of interest in consecutive frames is determined, including: Morphological dilation is performed on the region of interest in the images of consecutive video frames to cover the minute movement of the target in the region of interest, thereby expanding the space. A first-order low-pass filter or Kalman filter is applied to the center coordinates and size of the bounding box of the region of interest between consecutive video frames to smooth the motion trajectory and avoid box jumps, thereby achieving temporal smoothing.

5. The method according to any one of claims 1-3, characterized in that, The stabilized region of interest includes a primary region of interest and a secondary region of interest. Based on the stabilized region of interest, the real-time video stream is subjected to region-level coding to obtain coded video data. The image quality parameters of the stabilized region of interest are superior to those of the unstable region of interest, including: For the region of main interest, set it as the first image quality parameter; For the secondary region of interest, set it as the second image quality parameter; For the background region, set it to the third image quality parameter; Based on the first image quality parameter, the second image quality parameter, and the third image quality parameter; Among them, the values ​​of the first image quality parameter, the second image quality parameter, and the third image quality parameter increase sequentially.

6. The method according to any one of claims 1-3, characterized in that, The step of allocating the adaptive bitstream to the corresponding transmission path according to the quality level includes: When the current network bandwidth is within the first threshold range, the quality level of the adaptive bitstream is determined to be high, and the region of interest of the video image after stabilization is adopted in extended mode, and the first image quality parameter is adopted based on the region of interest after stabilization. When the current network bandwidth is in the second threshold range, the quality level of the adaptive bitstream is determined to be high, and the standard mode is adopted for the region of interest of the video image after stabilization processing, and the second image quality parameter is adopted for the region of interest after stabilization processing. When the current network bandwidth is in the third threshold range, the quality level of the adaptive bitstream is determined to be high, and the core mode is adopted for the region of interest of the video image after stabilization processing, and the third image quality parameter is adopted for the region of interest after stabilization processing.

7. A global live streaming bandwidth saving device based on region of interest optimization, characterized in that, The device includes: The recognition module is used to perform region of interest analysis on the received real-time video stream and identify the region of interest in each frame of the real-time video stream. The stabilization module is used to stabilize the region of interest in each frame of video image to obtain a stabilized region of interest; the stabilization process is used to spatially expand and temporally smooth the boundary of the region of interest. The encoding module is used to perform regional hierarchical encoding on the real-time video stream based on the stabilized region of interest to obtain hierarchically encoded video data; wherein the image quality parameters of the stabilized region of interest are better than those of the unstable region of interest. The encapsulation module is used to encapsulate hierarchically coded video data into an adaptive bitstream containing multiple quality levels; wherein, in an adaptive bitstream with a higher quality level, the area of ​​the region of interest based on the stabilized processing of the video image is larger or the image quality parameter of the video image is larger. The allocation module is used to allocate the adaptive bitstream to the corresponding transmission path according to the quality level.

8. The apparatus according to claim 7, characterized in that, The step of performing region of interest (ROI) analysis on the received real-time video stream to identify the ROI in each frame of the video stream further includes: Multimodal region of interest detection is performed on the video image to obtain visual detection results; Receive the marking information of the focal region coordinates in the video image from the broadcast control system; Receive user interaction behavior statistics; Based on the visual detection results, the labeling information, and the interaction behavior statistics, the region of interest in the video image and the detection confidence of the region of interest in the video image are determined.

9. A computer device, characterized in that, include: The processor, memory, communication interface, and communication bus are provided, wherein the processor, memory, and communication interface communicate with each other via the communication bus. The memory is used to store at least one executable instruction that causes the processor to perform the operation of the global live streaming bandwidth saving method based on region of interest optimization as described in any one of claims 1-8.

10. A computer-readable storage medium, characterized in that, The storage medium stores at least one executable instruction, which, when executed on a computer device, causes the computer device to perform the operation of the global live streaming bandwidth saving method based on region of interest optimization as described in any one of claims 1-8.