A method and related device for dynamic adjustment of video stream code rate
Patent Information
- Application Number
- CN202611062369.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-16
- Publication Date
- 2026-09-29
AI Technical Summary
[0003]本发明实施例的主要目的在于提出一种视频流码率动态调整方法、装置、电子设备、存储介质及程序产品,旨在解决现有技术的至少一种问题
[0015]本发明实施例至少包括以下有益效果:本发明提供一种视频流码率动态调整方法、装置、电子设备、存储介质及程序产品,该方案通过获取实时视频流中的原始视频帧序列;对原始视频帧序列中的视频帧进行分级分析,生成内容特征向量;其中,内容特征向量包括场景语义特征、运动强度特征和关注区域特征;基于内容特征向量和预设的目标带宽上限,通过参数决策生成动态压缩参数集;利用动态压缩参数集进行码率调整和推流。本发明通过对视频帧进行分级分析,能够提取场景语义、运动强度和关注区域等多维内容特征以生成结构化的内容特征向量,进而可以准确反映当前画面的复杂度与信息重要性分布;基于内容特征向量结合预设目标带宽上限进行参数决策,能够动态生成与内容特性及带宽约束精确匹配的压缩参数集,以实现码率的精细化调整;基于此,本发明可以在复杂场景或关注区域自动提升码率以保证关键信息清晰度,在简单或背景区域相应降低码率以节省带宽,能够克服传统固定码率或纯网络自适应方案无法兼顾内容差异的缺陷,可以显著提高带宽利用效率和视频整体视觉质量。
Smart Images

Figure CN122845875A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image processing technology, and in particular to a method and related equipment for dynamically adjusting the bitrate of a video stream. Background Technology
[0002] In the field of real-time video streaming, bitrate adjustment is a key technology for balancing image quality and transmission stability. Currently, mainstream bitrate control schemes are mainly divided into two categories: Fixed Bitrate (CBR) and Adaptive Bitrate Based on Network Conditions (ABR). The CBR scheme maintains a constant encoding bitrate throughout the transmission process, completely ignoring real-time changes in video content complexity and network bandwidth fluctuations. When the scene features high-speed motion or rich textures, the fixed bitrate is often insufficient to guarantee clarity, leading to blurry images or blockiness. Conversely, in static and simple scenes, the fixed bitrate wastes bandwidth resources. Adaptive Bitrate Based on Network Conditions dynamically switches video streams with different bitrates by monitoring channel parameters such as packet loss rate, latency, and throughput, offering some adaptability to fluctuations in network conditions. However, its decision-making is limited to the immediate conditions of the transmission channel and does not perceive differences in the video content itself, thus failing to perform fine-grained bitrate adjustments based on changes in content complexity. Furthermore, some existing schemes attempt to utilize low-level statistical features of video frames, such as motion vectors and texture complexity, to assist bitrate decisions, but these methods only address low-level information at the pixel level. Summary of the Invention
[0003] The main objective of this invention is to provide a method, apparatus, electronic device, storage medium, and program product for dynamic adjustment of video stream bitrate, aiming to solve at least one problem in the prior art.
[0004] To achieve the above objectives, one aspect of this invention proposes a method for dynamically adjusting the bitrate of a video stream, the method comprising: Obtain the raw video frame sequence from the real-time video stream; The video frames in the original video frame sequence are subjected to hierarchical analysis to generate content feature vectors; the content feature vectors include scene semantic features, motion intensity features, and region of interest features. Based on content feature vectors and preset target bandwidth limits, a dynamic compression parameter set is generated through parameter decision-making. Bitrate adjustment and streaming are performed using a dynamic compression parameter set.
[0005] In some embodiments, hierarchical analysis of video frames in the original video frame sequence includes the following steps: Inter-frame difference detection is performed on each adjacent video frame in the original video frame sequence to obtain the inter-frame difference value of each adjacent video frame. If the inter-frame difference value is greater than the difference threshold, the next video frame of the corresponding adjacent video frame is marked as a key frame. The pre-defined AI model is invoked to divide the original video frame sequence into several image groups, and the video frame corresponding to the starting point of each image group is marked as a keyframe. Iterate through each video frame in the original video frame sequence. If the video frame is a keyframe, generate a content feature vector through depth analysis; otherwise, generate a content feature vector through lightweight tracing.
[0006] In some embodiments, generating content feature vectors through deep analysis includes the following steps: Using a large AI model to perform scene semantic understanding on video frames, scene semantic features are obtained; these features include scene labels and semantic segmentation maps. The motion complexity of video frames is analyzed using a large AI model to obtain motion intensity features, which include global motion intensity indices and local motion vector fields. The AI large model is used to identify regions of interest in video frames to obtain features of the regions of interest; these features include the location coordinates and priority level of the regions of interest. The scene semantic features, motion intensity features, and region of interest features are summarized and encoded to generate the content feature vector of the keyframe.
[0007] In some embodiments, generating content feature vectors via lightweight tracking includes the following steps: Use the content feature vector of the previous key frame in the original video frame sequence as the reference feature vector for the current non-key frame video frame. The position coordinates of the reference feature vector are incrementally updated in the current non-key frame video frame using an optical flow tracker. The position coordinates in the baseline feature vector are replaced with the incrementally updated position coordinates to generate the content feature vector of the current non-key frame.
[0008] In some embodiments, a dynamic compression parameter set is generated based on the content feature vector and a preset target bandwidth limit through parameter decision-making, including the following steps: Based on the content feature vector and the preset target bandwidth limit, the dynamic compression parameter set is determined by querying the predefined decision rule base mapping. Alternatively, the target bandwidth limit can be used as a constraint, and the optimization problem of maximizing the estimated video quality function under the pre-modeled constraints can be solved using the content feature vector to obtain a dynamic compression parameter set. The dynamic compression parameter set includes bitrate, resolution, frame rate, and image group size.
[0009] In some embodiments, the content feature vector includes the location coordinates of the region of interest and its priority level. When the priority level is a first-level region of interest, the method further includes the following steps: The bitrate of the dynamic compression parameter set is increased to obtain the regional bitrate; The regional bitrate is associated with the location coordinates corresponding to the primary region of interest and added to the dynamic compression parameter set.
[0010] In some embodiments, bitrate adjustment and streaming are performed using a dynamic compression parameter set, including the following steps: Obtain the dynamic compression parameter set of the current video frame, and then use the dynamic compression parameter set to encode the current video frame to obtain the compressed video frame corresponding to the current video frame. The compressed video frames are pushed to the server, and the next video frame is subjected to hierarchical analysis and parameter decision-making in an asynchronous pipeline manner to generate a dynamic compression parameter set for the next video frame.
[0011] To achieve the above objectives, another aspect of the present invention provides a video stream bitrate dynamic adjustment device, the device comprising: The first module is used to acquire the original video frame sequence in the real-time video stream; The second module is used to perform hierarchical analysis on video frames in the original video frame sequence and generate content feature vectors; the content feature vectors include scene semantic features, motion intensity features, and region of interest features. The third module is used to generate a dynamic compression parameter set based on the content feature vector and the preset target bandwidth limit through parameter decision-making. The fourth module is used for bitrate adjustment and streaming using a dynamic compression parameter set.
[0012] To achieve the above objectives, another aspect of the present invention provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the aforementioned method.
[0013] To achieve the above objectives, another aspect of the present invention provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the aforementioned method.
[0014] To achieve the above objectives, another aspect of the present invention provides a computer program product, including a computer program that, when executed by a processor, implements the aforementioned method.
[0015] The embodiments of the present invention include at least the following beneficial effects: The present invention provides a method, apparatus, electronic device, storage medium, and program product for dynamic adjustment of video stream bitrate. This solution obtains the original video frame sequence in a real-time video stream; performs hierarchical analysis on the video frames in the original video frame sequence to generate a content feature vector; wherein, the content feature vector includes scene semantic features, motion intensity features, and region of interest features; based on the content feature vector and a preset target bandwidth limit, a dynamic compression parameter set is generated through parameter decision; and the bitrate is adjusted and the stream is pushed using the dynamic compression parameter set. This invention performs hierarchical analysis on video frames to extract multi-dimensional content features such as scene semantics, motion intensity, and regions of interest to generate structured content feature vectors. This accurately reflects the complexity and information importance distribution of the current scene. Based on the content feature vectors and a preset target bandwidth limit, parameter decisions are made to dynamically generate a set of compression parameters that precisely matches the content characteristics and bandwidth constraints, enabling fine-tuning of the bitrate. Therefore, this invention can automatically increase the bitrate in complex scenes or regions of interest to ensure the clarity of key information, and correspondingly decrease the bitrate in simple or background areas to save bandwidth. This overcomes the shortcomings of traditional fixed bitrate or pure network adaptive solutions that cannot accommodate content differences, significantly improving bandwidth utilization efficiency and overall video visual quality. Attached Figure Description
[0016] Figure 1 This is a schematic diagram of an implementation environment for the method of dynamically adjusting the bitrate of a video stream provided in an embodiment of the present invention; Figure 2 This is a flowchart illustrating the video stream bitrate dynamic adjustment method provided in an embodiment of the present invention; Figure 3 This is a schematic diagram of the overall architecture of the video stream bitrate dynamic adjustment method provided in this embodiment of the invention; Figure 4 This is a schematic diagram illustrating the overall process of the video stream bitrate dynamic adjustment method provided in this embodiment of the invention; Figure 5 This is a schematic diagram illustrating an example of an asynchronous processing flow provided in an embodiment of the present invention; Figure 6 This is a schematic diagram of the structure of the video stream bitrate dynamic adjustment device provided in an embodiment of the present invention; Figure 7 This is a schematic diagram of the structure of the electronic device provided in an embodiment of the present invention. Detailed Implementation
[0017] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention. In the following description, when referring to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the embodiments of this invention; they are merely examples of apparatuses and methods consistent with some aspects of the embodiments of this invention as detailed in the appended claims.
[0018] It is understood that the terms "first," "second," etc., used in this invention may be used to describe various concepts, but unless specifically stated otherwise, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another. For example, without departing from the scope of embodiments of this invention, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the words "if" or "when" as used herein may be interpreted as "when," "in response to determination," or "in the event of a determination."
[0019] The terms “at least one,” “multiple,” “each,” “any,” etc., used in this invention, “at least one” includes one, two, or more than two; “multiple” includes two or more than two; “each” refers to each of the corresponding multiple; and “any” refers to any one of the multiple.
[0020] Unless otherwise defined, all technical and scientific terms used in this invention have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. The terminology used in this invention is for descriptive purposes only and is not intended to limit the invention.
[0021] To facilitate understanding of the technical solution of this invention, the following explanations are provided regarding the technical terms that may be involved in the technical solution of this invention: Video streaming: The process of encapsulating and transmitting real-time generated video data to a streaming media server.
[0022] Large AI models refer to machine learning models capable of semantic understanding, motion analysis, and / or region of interest identification from video frames. These models can be pre-trained large models optimized through pruning, quantization, or highly efficient neural networks specifically trained to perform particular analytical tasks.
[0023] Compression parameters: These refer to the specific settings used to control compression efficiency and quality during video encoding, including but not limited to: encoder type (such as H.264, H.265), bitrate, frame rate, resolution, GOP (group of pictures) size, quantization parameters (QP), etc.
[0024] Streaming module: A software or hardware functional unit responsible for performing video encoding and network transmission tasks.
[0025] In related technologies, the lack of understanding of the overall semantic content of the video makes it difficult to identify targets of significant interest and to achieve differentiated coding based on content importance. These problems are particularly pronounced in scenarios such as security monitoring where uplink bandwidth resources are limited and multiple concurrent transmissions are occurring. How to ensure both real-time and stable transmission of the video stream under bandwidth constraints, while prioritizing the clarity of critical areas such as faces and license plates, remains an unresolved issue.
[0026] In view of this, this invention provides a method and related device for dynamic adjustment of video stream bitrate. The solution obtains the original video frame sequence in the real-time video stream; performs hierarchical analysis on the video frames in the original video frame sequence to generate content feature vectors; wherein, the content feature vectors include scene semantic features, motion intensity features, and region of interest features; based on the content feature vectors and a preset target bandwidth limit, a dynamic compression parameter set is generated through parameter decision; and the bitrate is adjusted and the stream is pushed using the dynamic compression parameter set. This invention performs hierarchical analysis on video frames to extract multi-dimensional content features such as scene semantics, motion intensity, and regions of interest to generate structured content feature vectors. This accurately reflects the complexity and information importance distribution of the current scene. Based on the content feature vectors and a preset target bandwidth limit, parameter decisions are made to dynamically generate a set of compression parameters that precisely matches the content characteristics and bandwidth constraints, enabling fine-tuning of the bitrate. Therefore, this invention can automatically increase the bitrate in complex scenes or regions of interest to ensure the clarity of key information, and correspondingly decrease the bitrate in simple or background areas to save bandwidth. This overcomes the shortcomings of traditional fixed bitrate or pure network adaptive solutions that cannot accommodate content differences, significantly improving bandwidth utilization efficiency and overall video visual quality.
[0027] It is understood that the video stream bitrate dynamic adjustment method provided by this invention can be applied to any computer device with data processing and computing capabilities, and this computer device can be various terminals or servers. When the computer device in the embodiment is a server, the server is an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. Optionally, the terminal can be a smartphone, tablet, laptop, or desktop computer, but it is not limited to these.
[0028] like Figure 1 The diagram shown is a schematic representation of an implementation environment provided by an embodiment of the present invention. (Refer to...) Figure 1 The implementation environment includes at least one terminal 102 and a server 101. The terminal 102 and the server 101 can be connected via a network, either wirelessly or via a wired connection, to complete data transmission and exchange.
[0029] Server 101 can be a standalone physical server, a server cluster or distributed system consisting of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms.
[0030] Additionally, server 101 can also be a node server in a blockchain network. Blockchain is a novel application model of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanisms, and encryption algorithms.
[0031] Terminal 102 can be a smartphone, tablet computer, laptop computer, desktop computer, smart speaker, smartwatch, etc., but is not limited to these. Terminal 102 and server 101 can be directly or indirectly connected via wired or wireless communication, and this embodiment of the invention does not impose any limitations.
[0032] For example, based on Figure 1 The implementation environment shown in this embodiment of the invention provides a method for dynamically adjusting the bitrate of a video stream. The following description uses the application of this method in server 101 as an example. It can be understood that this method can also be applied in terminal 102.
[0033] Reference Figure 2 , Figure 2 This is an optional flowchart of the video stream bitrate dynamic adjustment method provided in the embodiments of the present invention. The execution subject of the video stream bitrate dynamic adjustment method can be any of the aforementioned computer devices (including servers or terminals). Figure 2 The method may include, but is not limited to, steps S100 to S400.
[0034] Step S100: Obtain the original video frame sequence from the real-time video stream; For example, in some specific implementations, raw video frame sequences can be acquired from a real-time video stream using a pre-deployed video acquisition module.
[0035] Step S200: Perform hierarchical analysis on the video frames in the original video frame sequence to generate content feature vectors; The content feature vector includes scene semantic features, motion intensity features, and region of interest features; It should be noted that, in some embodiments, hierarchical analysis of video frames in the original video frame sequence may include the following steps: performing inter-frame difference detection on each adjacent video frame in the original video frame sequence to obtain the inter-frame difference value of each adjacent video frame; if the inter-frame difference value is greater than the difference threshold, marking the next video frame of the corresponding adjacent video frame as a key frame; calling a preset AI large model to divide the original video frame sequence into several image groups, and marking the video frame corresponding to the starting point of each image group as a key frame; traversing each video frame in the original video frame sequence, when the video frame is a key frame, generating a content feature vector through deep analysis; otherwise, generating a content feature vector through lightweight tracking.
[0036] For example, in some specific implementations, the video acquisition module inputs the raw video frame sequence to the AI analysis module. The module's hierarchical processing and scheduling unit first classifies the frames: when a drastic scene change is detected (through inter-frame difference detection) or a new GOP (Group of Pictures) starting point is reached, the frame is marked as a "keyframe" and sent to the "deep analysis" channel. Otherwise, the frame is marked as a "non-keyframe" and sent to the "lightweight tracking" channel.
[0037] It should be noted that, in some embodiments, generating content feature vectors through deep analysis may include the following steps: using a large AI model to perform scene semantic understanding on video frames to obtain scene semantic features; wherein, scene semantic features include scene labels and semantic segmentation maps; using a large AI model to perform motion complexity analysis on video frames to obtain motion intensity features; wherein, motion intensity features include global motion intensity indices and local motion vector fields; using a large AI model to identify regions of interest on video frames to obtain regions of interest features; wherein, regions of interest features include the location coordinates and priority level of the regions of interest; and summarizing and encoding the scene semantic features, motion intensity features, and regions of interest features to generate content feature vectors for keyframes.
[0038] For example, in some specific implementations, keyframe depth analysis can be implemented as follows: for keyframes, launch a full version of the large AI model (pruned and quantized for optimization) and perform the following analysis in parallel: a. Scene semantic understanding: Technical means: using the image classification or scene segmentation capabilities of large models.
[0039] Output characteristics: Scene tags: such as "indoor office", "outdoor street", "stadium", "natural scenery", etc. The average texture complexity varies significantly across different scenes.
[0040] Semantic segmentation map: This classifies elements in the image, such as the sky, buildings, faces, screen, and plants, at the pixel level. This provides the foundation for subsequent "Region of Interest (ROI)" identification.
[0041] b. Motion complexity analysis: Technical means: Utilize the temporal modeling capabilities of large models (such as 3D convolution, optical flow estimation layers) or combine them with classical computer vision methods.
[0042] Output characteristics: Global Motion Intensity Index: A value between 0 and 1, representing the intensity of motion of objects in the entire image. For example, a pedestrian walking is approximately 0.1, a non-motorized vehicle moving is 0.4, while a car moving at high speed can reach 0.7 or higher.
[0043] Local motion vector field: describes the direction and speed of motion of each region or object in the image.
[0044] c. Region of Interest (ROI) Identification: Technical means: Combining the results of the above semantic segmentation and object detection (such as YOLO, Faster R-CNN, etc., which can be integrated into a large model).
[0045] Output characteristics: ROI Location and Ranking: Identify the location (such as a rectangle or mask) of one or more key areas in the image and assign them an importance ranking. For example: First-level ROI (highest priority): face area, license plate area.
[0046] Secondary ROI (medium priority): Motor vehicle location, non-motor vehicle location, human location.
[0047] Background (low priority): background objects such as walls, sky, and buildings.
[0048] It should be noted that, in some embodiments, generating content feature vectors through lightweight tracking may include the following steps: taking the content feature vector corresponding to the previous key frame in the original video frame sequence as the reference feature vector of the current non-key frame video frame; using an optical flow tracker to incrementally update the position coordinates of the reference feature vector in the current non-key frame video frame; and replacing the position coordinates in the reference feature vector with the incrementally updated position coordinates to generate the content feature vector of the current non-key frame.
[0049] For example, in some specific implementations, lightweight tracking of non-keyframes can be achieved as follows: for non-keyframes, instead of starting a large model, a lightweight tracker with minimal computational cost (such as one based on optical flow) is used to fine-tune and update only the analysis results of the previous keyframe, such as updating the position of the ROI.
[0050] Step S300: Based on the content feature vector and the preset target bandwidth limit, a dynamic compression parameter set is generated through parameter decision-making; It should be noted that in some embodiments, the generation of a dynamic compression parameter set based on content feature vectors and a preset target bandwidth limit through parameter decision may include the following steps: determining the dynamic compression parameter set by querying a predefined decision rule base mapping based on content feature vectors and a preset target bandwidth limit; or, using the target bandwidth limit as a constraint, using the content feature vectors to solve the optimization problem of maximizing the estimated video quality function under pre-modeled constraints to obtain the dynamic compression parameter set; wherein, the dynamic compression parameter set includes bitrate, resolution, frame rate, and image group size.
[0051] For example, in some specific implementations, the parameter decision module is configured to map or calculate a set of compression parameters based on the content feature vector and the target bandwidth constraint, using a predefined strategy or optimization algorithm. The decision module receives the "content feature vector" and the system-preset "target bandwidth limit".
[0052] It outputs the final "compression parameter set" through one or more of the following mechanisms: Decision-making mechanism one (based on predefined rules / lookup tables): How it works: It internally maintains a large base of decision rules or lookup tables.
[0053] Example rules: IF Scene == Static Scene AND Motion Intensity < 0.2 THEN Baseline Bitrate = 500kbps.
[0054] IF Scene == Dynamic Scene AND Motion Intensity > 0.7 THEN Baseline Bitrate = 2000kbps.
[0055] If global motion intensity > 0.5THEN, reduce the GOP size from 250 to 60 to reduce motion prediction error.
[0056] Decision-making mechanism two (based on a lightweight optimization model): How it works: The decision problem is modeled as an optimization problem: maximizing the "estimated video quality" under the constraint of "target bandwidth".
[0057] Objective function: Maximize: Q = f(content feature vector, compression parameters); Constraint: Output_Bitrate (compression parameter) ≤ Target_Bandwidth; Execution process: The module uses a lightweight machine learning model (such as a gradient boosting decision tree or a small neural network) to quickly predict the output bitrate and quality score under different parameter combinations, and selects the parameter set with the highest score.
[0058] Output the compression parameter set and execute: The decision module outputs a specific set of parameters that can be directly invoked by the streaming encoding module. These parameters are dynamic and change with the content.
[0059] It should be noted that the content feature vector includes the location coordinates and priority level of the region of interest. In some embodiments, when the priority level is a first-level region of interest, the method may also include the following steps: boosting the bitrate in the dynamic compression parameter set to obtain the region bitrate; associating the region bitrate with the location coordinates corresponding to the first-level region of interest and supplementing it to the dynamic compression parameter set.
[0060] For example, in some specific implementations, a rule can also be configured: if a first-level ROITHEN exists, ROI encoding is enabled, and the bitrate for that region is increased by 50%.
[0061] Step S400: Use the dynamic compression parameter set to adjust the bitrate and push the stream; It should be noted that in some embodiments, bitrate adjustment and streaming using a dynamic compression parameter set may include the following steps: obtaining the dynamic compression parameter set of the current video frame, then encoding the current video frame using the dynamic compression parameter set to obtain the compressed video frame corresponding to the current video frame; streaming the compressed video frame to the server, and performing hierarchical analysis and parameter decision-making on the next video frame in an asynchronous pipeline manner to generate a dynamic compression parameter set for the next video frame.
[0062] For example, in some specific embodiments, this set of dynamic parameters can be received by the push-stream encoding module and immediately applied to the current video stream encoder. The compressed video stream is then pushed to the server. In some preferred embodiments, the technical solution of this invention employs asynchronous pipelined operation: the AI analysis module and the push-stream encoding module work in parallel. While the Nth frame is being encoded, the AI analysis module is already analyzing the N+1th frame in parallel. This design "hides" most of the analysis time within the encoding time, greatly eliminating perceived latency at the system level.
[0063] To explain in detail the principle of the technical solution of the present invention, the overall process of the present invention will be described below with reference to some specific embodiments. It is easy to understand that the following is an explanation of the technical principle of the present invention and should not be regarded as a limitation of the present invention.
[0064] First, it's important to note that existing technologies suffer from the following main problems: With limited server bandwidth, they cannot dynamically and accurately allocate encoding resources based on the real-time complexity of the video content. This leads to either wasted bandwidth (high bitrate for simple content) or severely compromised video quality (low bitrate for complex content, resulting in pixelation, blurring, etc.). Adaptive strategies suffer from high response latency and cannot adapt to rapid content changes; their simplistic approach fails to comprehensively balance content value (e.g., high fidelity is required for faces and text regions) with overall bandwidth constraints. Furthermore, existing solutions lack deep, real-time semantic understanding of the video stream content before or during encoding. Therefore, they cannot predict changes in the complexity of subsequent frames and can only passively and reactively adjust the bitrate, which is the root cause of the aforementioned technical problems.
[0065] In view of the shortcomings of existing technologies, this invention introduces a large AI model to analyze the input video stream in real time, predict the best compression parameters that balance compression efficiency and visual quality, and dynamically feeds these parameters back to the streaming module to guide it in performing accurate compression.
[0066] like Figure 3 The diagram shows an example of the architecture of the technical solution of this invention. In some specific application scenarios, the system architecture of the technical solution of this invention may include a video acquisition module, an AI analysis module, a parameter decision module, and a streaming encoding module. Its data flow and closed-loop control logic are as follows: Video capture module: responsible for acquiring the original video frame sequence.
[0067] AI Analysis Module: Connects to the video acquisition module to receive video frames. This module incorporates an optimized and lightweight AI model for performing hierarchical content analysis tasks.
[0068] Parameter decision module: Connects to the AI analysis module, receives the "content feature vector" output by the AI analysis module, and calculates the optimal set of compression parameters based on preset bandwidth limitation rules or optimization algorithms.
[0069] Streaming encoding module: Connected to the parameter decision module, it receives the calculated dynamic compression parameters and uses them to encode and stream the video stream in real time.
[0070] like Figure 4 The diagram shows an overall process example of the technical solution of this invention. In some specific applications, it can be implemented through the following process: Step 1: Video Frame Input and Scheduling The video acquisition module inputs the raw video frame sequence into the AI analysis module. The module's hierarchical processing and scheduling unit first classifies the frames: when a drastic change in the scene is detected (through inter-frame difference detection) or a new GOP (Group of Pictures) start point is reached, the frame is marked as a "keyframe" and sent to the "deep analysis" channel.
[0071] Otherwise, mark the frame as a "non-critical frame" and send it to the "lightweight tracking" channel.
[0072] Step 2: Hierarchical Analysis of Large AI Models (1) Keyframe in-depth analysis: For keyframes, start the full version of the large AI model (after pruning and quantization optimization) and perform the following analysis in parallel: a. Scene semantic understanding: Technical means: using the image classification or scene segmentation capabilities of large models.
[0073] Output characteristics: Scene tags: such as "indoor office", "outdoor street", "stadium", "natural scenery", etc. The average texture complexity varies significantly across different scenes.
[0074] Semantic segmentation map: This classifies elements in the image, such as the sky, buildings, faces, screen, and plants, at the pixel level. This provides the foundation for subsequent "Region of Interest (ROI)" identification.
[0075] b. Motion complexity analysis: Technical means: Utilize the temporal modeling capabilities of large models (such as 3D convolution, optical flow estimation layers) or combine them with classical computer vision methods.
[0076] Output characteristics: Global Motion Intensity Index: A value between 0 and 1, representing the intensity of motion of objects in the entire image. For example, a pedestrian walking is approximately 0.1, a non-motorized vehicle moving is 0.4, while a car moving at high speed can reach 0.7 or higher.
[0077] Local motion vector field: describes the direction and speed of motion of each region or object in the image.
[0078] c. Region of Interest (ROI) Identification: Technical means: Combining the results of the above semantic segmentation and object detection (such as YOLO, Faster R-CNN, etc., which can be integrated into a large model).
[0079] Output characteristics: ROI Location and Ranking: Identify the location (such as a rectangle or mask) of one or more key areas in the image and assign them an importance ranking. For example: First-level ROI (highest priority): face area, license plate area.
[0080] Secondary ROI (medium priority): Motor vehicle location, non-motor vehicle location, human location.
[0081] Background (low priority): background objects such as walls, sky, and buildings.
[0082] (2) Lightweight tracking of non-keyframes: For non-keyframes, instead of starting a large model, a lightweight tracker with minimal computational cost (such as one based on optical flow) is used to fine-tune and update only the analysis results of the previous keyframe, such as updating the position of the ROI.
[0083] (3) Generate content feature vectors: The AI analysis module summarizes and encodes all the above analysis results (scene labels, motion intensity, ROI location and level, etc.) to form a structured, low-dimensional "content feature vector".
[0084] Optional example vector: Forms a structured 'content feature vector'. For example, the vector may include one or more combinations of the following features: scene encoding, global motion intensity, number, type, location coordinates, area percentage, and priority of regions of interest (ROIs).
[0085] For example: [Scene encoding=12, Global motion intensity=0.85, Number of ROIs=2, ROI1_Type=Face, ROI1_Center coordinates=(0.6,0.4), ROI1_Area percentage=0.1, ROI1_Priority=1, ROI2_Type=Vehicle, ROI2_Center coordinates=(0.5,0.7), ROI2_Area percentage=0.15, ROI2_Priority=1]; This vector is a digital summary of the current video content, which will serve as input to the parameter decision module.
[0086] Step 3: The parameter decision module performs mapping and optimization. The parameter decision module is configured to map or calculate a set of compression parameters based on the content feature vector and the target bandwidth constraint, using a predefined strategy or optimization algorithm. The decision module receives the "content feature vector" and the system-preset "target bandwidth limit".
[0087] It outputs the final "compression parameter set" through one or more of the following mechanisms: Decision-making mechanism one (based on predefined rules / lookup tables): How it works: It internally maintains a large base of decision rules or lookup tables.
[0088] Example rules: IF Scene == Static Scene AND Motion Intensity < 0.2 THEN Baseline Bitrate = 500kbps.
[0089] IF Scene == Dynamic Scene AND Motion Intensity > 0.7 THEN Baseline Bitrate = 2000kbps.
[0090] If a first-level ROUTHEN exists, ROI encoding is enabled, increasing the bitrate of that region by 50%.
[0091] If global motion intensity > 0.5THEN, reduce the GOP size from 250 to 60 to reduce motion prediction error.
[0092] Decision-making mechanism two (based on a lightweight optimization model): How it works: The decision problem is modeled as an optimization problem: maximizing the "estimated video quality" under the constraint of "target bandwidth".
[0093] Objective function: Maximize: Q = f(content feature vector, compression parameters); Constraint: Output_Bitrate (compression parameter) ≤ Target_Bandwidth; Execution process: The module uses a lightweight machine learning model (such as a gradient boosting decision tree or a small neural network) to quickly predict the output bitrate and quality score under different parameter combinations, and selects the parameter set with the highest score.
[0094] Output the compression parameter set and execute: The decision module outputs a specific set of parameters that can be directly invoked by the streaming encoding module. These parameters are dynamic and change with the content.
[0095] Example output 1 (static scene): Content: Pedestrians wait at the traffic light on the side of the road.
[0096] AI analysis results: Scene = Outdoor, Motion intensity = 0.2, 1 face ROI detected.
[0097] Decision output parameters: { "codec": "H.265", "resolution": "720p", "bitrate":600kbps, "framerate": 15, "gop_size": 120, "roi_enabled": true, "roi_list":[{"coordinates": [...], "boost": 1.5}]}.
[0098] Example output 2 (dynamic scene): Content: A car traveling at high speed.
[0099] AI analysis results: Scene = Street, Motion Intensity = 0.9, 3 vehicle ROIs detected.
[0100] Decision output parameters: { "codec": "H.265", "resolution": "1080p", "bitrate":2500kbps, "framerate": 30, "gop_size": 45, "roi_enabled": false}.
[0101] Step 4: Encoding and Streaming The streaming encoding module receives this set of dynamic parameters and immediately applies them to the current video stream encoder. The compressed video stream is then pushed to the server.
[0102] like Figure 5 As shown, this is an example of an asynchronous processing flow provided by an embodiment of the present invention. In some specific application scenarios, the technical solution of the present invention also implements the following real-time guarantee mechanism: Asynchronous pipelined operation: The AI analysis module and the push-stream encoding module work in parallel. As shown in the diagram above, while frame N is being encoded, the AI analysis module is already analyzing frame N+1 in parallel. This design "hides" most of the analysis time within the encoding time, greatly eliminating perceived latency at the system level.
[0103] Hardware acceleration: The AI analysis module is deployed on dedicated AI inference hardware (such as GPU / NPU) to ensure single-frame analysis speed in the millisecond range.
[0104] Analysis result inheritance: The results of a deep analysis will remain effective for a period of time (such as one GOP cycle), avoiding frequent parameter switching and ensuring system stability.
[0105] In summary, the technical solution of this invention constructs a content-aware adaptive encoding system, enabling a leap from "pipeline awareness" to "content awareness." Specifically, it performs deep analysis on keyframes, including scene semantic understanding, motion complexity analysis, and region of interest identification; and for non-keyframes, it uses a lightweight tracker to incrementally update the results of the deep analysis. This significantly reduces computational overhead while maintaining analysis accuracy, ensuring the system's real-time performance. Furthermore, the parameter decision module determines the compression parameter set by querying a predefined strategy library or running optimization algorithms, providing a flexible and efficient decision-making mechanism that can finely balance bandwidth limitations and visual quality. Simultaneously, the AI analysis module and the streaming encoding module work in parallel in an asynchronous pipeline manner, hiding AI analysis time within the encoding time, eliminating analysis latency at the system level, and ensuring real-time transmission of the video stream.
[0106] Compared with the prior art, the present invention has at least the following beneficial effects: 1. Intelligent: It can understand video content and allocate bitrate as needed to maximize bandwidth utilization.
[0107] 2. High real-time performance: Through hierarchical analysis and asynchronous pipeline design, ultra-low latency is ensured to meet real-time business requirements.
[0108] 3. Technological Integration: The computationally intensive large AI model was successfully combined with the latency-sensitive real-time video coding system, resolving the technical contradictions in combining the two.
[0109] 4. Architectural Innovation: The design incorporates "hierarchical analysis" and "asynchronous pipeline" mechanisms, which are key to achieving a balance between intelligence and real-time performance.
[0110] 5. Significant results: Compared with traditional solutions, it can achieve a significantly improved subjective image quality experience under the same bandwidth, demonstrating significant technological progress.
[0111] like Figure 6 As shown, this embodiment of the invention also provides a video stream bitrate dynamic adjustment device 900, which can implement the above-described method. This device may include: The first module 901 is used to acquire the original video frame sequence in the real-time video stream; The second module 902 is used to perform hierarchical analysis on video frames in the original video frame sequence and generate content feature vectors; wherein, the content feature vectors include scene semantic features, motion intensity features and region of interest features; The third module 903 is used to generate a dynamic compression parameter set based on the content feature vector and the preset target bandwidth limit through parameter decision-making. Module 4, 904, is used for bitrate adjustment and streaming using a dynamic compression parameter set.
[0112] It is understood that the content of the above method embodiments is applicable to the present device embodiments. The specific functions implemented by the present device embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0113] This invention also provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the method described above. This electronic device can be any smart terminal, including tablet computers, in-vehicle computers, etc.
[0114] It is understood that the content of the above method embodiments is applicable to this device embodiment. The specific functions implemented by this device embodiment are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0115] like Figure 7 As shown, Figure 7 The hardware structure of an electronic device 1000 according to another embodiment is illustrated. The electronic device 1000 includes: The processor 1001 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (aSIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of the present invention. The memory 1002 can be implemented as a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RaM). The memory 1002 can store the operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 1002 and is called and executed by the processor 1001. Input / output interface 1003 is used to implement information input and output; The communication interface 1004 is used to enable communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.). Bus 1005 transmits information between various components of the device (e.g., processor 1001, memory 1002, input / output interface 1003, and communication interface 1004); The processor 1001, memory 1002, input / output interface 1003 and communication interface 1004 are connected to each other within the device via bus 1005.
[0116] The electronic device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0117] This invention also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described method.
[0118] It is understood that the content of the above method embodiments is applicable to this storage medium embodiment. The specific functions implemented in this storage medium embodiment are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those achieved in the above method embodiments.
[0119] This invention also provides a computer program product, including a computer program that, when executed by a processor, implements the above-described method.
[0120] It is understood that the content of the above method embodiments is applicable to the embodiments of this program product. The specific functions implemented by the embodiments of this program product are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0121] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. Furthermore, memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, memory may optionally include memory remotely located relative to the processor, and these remote memories can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.
[0122] The video stream bitrate dynamic adjustment method, apparatus, electronic device, storage medium, and program product provided in this invention acquire the original video frame sequence in a real-time video stream; perform hierarchical analysis on the video frames in the original video frame sequence to generate a content feature vector; wherein the content feature vector includes scene semantic features, motion intensity features, and region of interest features; based on the content feature vector and a preset target bandwidth limit, a dynamic compression parameter set is generated through parameter decision; and the bitrate is adjusted and the stream is pushed using the dynamic compression parameter set. This invention performs hierarchical analysis on video frames to extract multi-dimensional content features such as scene semantics, motion intensity, and regions of interest to generate structured content feature vectors. This accurately reflects the complexity and information importance distribution of the current scene. Based on the content feature vectors and a preset target bandwidth limit, parameter decisions are made to dynamically generate a set of compression parameters that precisely matches the content characteristics and bandwidth constraints, enabling fine-tuning of the bitrate. Therefore, this invention can automatically increase the bitrate in complex scenes or regions of interest to ensure the clarity of key information, and correspondingly decrease the bitrate in simple or background areas to save bandwidth. This overcomes the shortcomings of traditional fixed bitrate or pure network adaptive solutions that cannot accommodate content differences, significantly improving bandwidth utilization efficiency and overall video visual quality.
[0123] The embodiments described in this invention are for the purpose of more clearly illustrating the technical solutions of the embodiments of this invention, and do not constitute a limitation on the technical solutions provided by the embodiments of this invention. As those skilled in the art will know, with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of this invention are also applicable to similar technical problems.
[0124] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present invention, and may include more or fewer steps than shown, or combine certain steps, or different steps.
[0125] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0126] Those skilled in the art will understand that all or some of the steps in the methods disclosed above, as well as the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, or suitable combinations thereof.
[0127] The preferred embodiments of the present invention have been described above with reference to the accompanying drawings, but this does not limit the scope of the claims of the present invention. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and spirit of the present invention should be within the scope of the claims of the present invention.
Claims
1. A method for dynamically adjusting the bitrate of a video stream, characterized in that, The method includes the following steps: Obtain the raw video frame sequence from the real-time video stream; The video frames in the original video frame sequence are subjected to hierarchical analysis to generate content feature vectors; wherein, the content feature vectors include scene semantic features, motion intensity features, and region of interest features; Based on the content feature vector and the preset target bandwidth limit, a dynamic compression parameter set is generated through parameter decision-making. The bitrate is adjusted and the stream is pushed using the dynamic compression parameter set.
2. The method according to claim 1, characterized in that, The hierarchical analysis of video frames in the original video frame sequence includes the following steps: Inter-frame difference detection is performed on each adjacent video frame in the original video frame sequence to obtain the inter-frame difference value of each adjacent video frame. If the inter-frame difference value is greater than the difference threshold, the next video frame of the corresponding adjacent video frame is marked as a key frame. The original video frame sequence is divided into several image groups by calling a preset AI large model, and the video frame corresponding to the starting point of each image group is marked as the key frame; Iterate through each video frame of the original video frame sequence. If the video frame is the keyframe, generate the content feature vector through depth analysis; otherwise, generate the content feature vector through lightweight tracking.
3. The method according to claim 2, characterized in that, The process of generating the content feature vector through deep analysis includes the following steps: The AI model is used to perform scene semantic understanding on video frames to obtain scene semantic features; wherein, the scene semantic features include scene labels and semantic segmentation maps; The AI large model is used to perform motion complexity analysis on video frames to obtain motion intensity features; wherein, the motion intensity features include global motion intensity index and local motion vector field; The AI model is used to identify regions of interest in video frames to obtain features of these regions; the features of these regions of interest include the location coordinates and priority level of the regions of interest. The scene semantic features, motion intensity features, and region of interest features are summarized and encoded to generate the content feature vector of the keyframe.
4. The method according to claim 3, characterized in that, The process of generating the content feature vector through lightweight tracking includes the following steps: The content feature vector corresponding to the previous key frame in the original video frame sequence of the current non-key frame is used as the reference feature vector. The position coordinates of the reference feature vector are incrementally updated in the current non-key frame video frame using an optical flow tracker. The position coordinates in the baseline feature vector are replaced with the incrementally updated position coordinates to generate the content feature vector of the current non-key frame.
5. The method according to claim 1, characterized in that, The process of generating a dynamic compression parameter set based on the content feature vector and a preset target bandwidth limit through parameter decision-making includes the following steps: Based on the content feature vector and the preset target bandwidth limit, the dynamic compression parameter set is determined by querying a predefined decision rule base mapping. Alternatively, the target bandwidth limit can be used as a constraint, and the content feature vector can be used to solve the optimization problem of maximizing the estimated video quality function under the pre-modeled constraint to obtain the dynamic compression parameter set; The dynamic compression parameter set includes bitrate, resolution, frame rate, and image group size.
6. The method according to claim 5, characterized in that, The content feature vector includes the location coordinates and priority level of the region of interest. When the priority level is a first-level region of interest, the method further includes the following steps: The bitrate of the dynamic compression parameter set is boosted to obtain the regional bitrate; The region bitrate is associated with the location coordinates corresponding to the first-level interest region and added to the dynamic compression parameter set.
7. The method according to claim 1, characterized in that, The process of adjusting bitrate and streaming using the dynamic compression parameter set includes the following steps: The dynamic compression parameter set of the current video frame is obtained, and then the current video frame is encoded using the dynamic compression parameter set to obtain the compressed video frame corresponding to the current video frame. The compressed video frame is pushed to the server, and the next video frame is subjected to the hierarchical analysis and parameter decision-making in an asynchronous pipeline manner to generate a dynamic compression parameter set for the next video frame.
8. A video stream bitrate dynamic adjustment device, characterized in that, The apparatus, applicable to the method according to any one of claims 1 to 7, comprises: The first module is used to acquire the original video frame sequence in the real-time video stream; The second module is used to perform hierarchical analysis on the video frames in the original video frame sequence and generate content feature vectors; wherein, the content feature vectors include scene semantic features, motion intensity features and region of interest features; The third module is used to generate a dynamic compression parameter set based on the content feature vector and the preset target bandwidth limit through parameter decision-making. The fourth module is used for bitrate adjustment and streaming using the dynamic compression parameter set.
9. An electronic device, characterized in that, The electronic device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the method according to any one of claims 1 to 7.
10. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the method according to any one of claims 1 to 7.