AVS video differential encoding method based on user attention mechanism
Patent Information
- Application Number
- CN202610805869.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-05
- Publication Date
- 2026-09-22
AI Technical Summary
[0006]本发明针对现有AVS3视频编码技术中码率分配缺乏用户意图导向、视频视觉区域映射精度不足、帧内帧间码率调度灵活性低、编码质量与码率利用率失衡的技术不足,提出一种基于用户注意力机制的AVS视频差异化编码方法,为一种基于注意力机制融合Grounding-DINO模型的差异化码率适配方法
本发明提出的基于注意力机制融合 Grounding-DINO 模型的差异化码率适配方法,深度贴合 AVS3 视频编码标准,针对性解决传统编码技术中码率分配无用户意图导向、文本与视觉区域映射精度不足、帧内帧间调度僵化、编码质量与码率利用率失衡的核心痛点。该方法以用户文本意图为核心,借助 Grounding-DINO 模型实现文本与视频视觉区域的精准锚定,结合注意力机制构建分层化码率分配体系,通过预设冗余码率暂存缓冲空间实现码率的动态回流复用,搭配闭环迭代校验机制优化编码效果,全程兼容 AVS3 编码规范,无需对现有编码框架进行大规模重构,部署成本低、适配场景广。其既保证了用户关注核心区域的编码质量,又最大化降低了非关键区域的码率消耗,显著提升码率资源利用率,为短视频、高清直播、智能监控等个性化视频编码场景提供了高效解决方案,具备极强的工程应用价值与技术创新性。
Smart Images

Figure CN122802679A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of video coding technology, specifically to an AVS video differential coding method based on user attention mechanism, which is applicable to real-time coding scenarios with limited bandwidth where priority must be given to ensuring the visual quality of user-focused areas. Background Technology
[0002] With the widespread adoption of ultra-high-definition video, real-time streaming media, and security monitoring, the amount of video data is growing exponentially. However, the growth rate of network transmission bandwidth and storage resources is relatively lagging behind, creating a core contradiction between video application demand and resource supply. AVS3, as a high-efficiency video coding standard independently developed in my country, has been widely used in various scenarios due to its excellent compression performance. However, its traditional global bitrate allocation logic has limitations in adaptability in specific scenarios where priority must be given to ensuring the visual quality of user-specific areas of interest.
[0003] Traditional AVS3 encoding, based on rate-distortion theory, performs global bitrate allocation and employs equalization logic, encoding all regions within a frame and all sequences between frames indiscriminately. This makes it difficult to respond to users' personalized semantic intent. For example, in a video scene containing a red sports car turning, traditional AVS3 encoding will evenly distribute the bitrate to the core frame of the turn, the preceding and following background frames, as well as the sports car's main body area and the background environment area. This results in valuable bitrate resources being wasted on non-interesting areas, and the subjective quality of the scenes and areas that users are most concerned about cannot be prioritized. This problem is particularly prominent in low-bandwidth, highly personalized scenarios.
[0004] While existing perceptual coding technologies attempt to incorporate visual saliency or simple semantic information to optimize bitrate allocation and compensate for the shortcomings of traditional coding, they still have significant drawbacks: First, bitrate scheduling has a single dimension, with most schemes focusing only on bitrate adjustment in the intra-frame spatial dimension, failing to explore the redundancy potential in the inter-frame temporal dimension, and unable to reuse bitrate saved from irrelevant frames in core frames, resulting in insufficient bitrate utilization; second, attention customization capabilities are insufficient, relying on fixed saliency models, unable to accurately respond to dynamic user intent through natural language, and exhibiting poor flexibility in adapting to different scenarios; third, there is a lack of closed-loop optimization mechanisms, relying solely on objective indicators to evaluate the effect after encoding, unable to incorporate user subjective feedback to correct strategies, and struggling to ensure that the encoding effect continuously meets actual needs.
[0005] Therefore, there is an urgent need for an encoding method that conforms to the AVS3 standard, can respond to user semantic intent, realizes intra-frame and inter-frame two-dimensional bitrate scheduling, and has closed-loop verification capability. This method should specifically address the problems of rigid bitrate allocation, insufficient intent response, and low resource utilization in traditional AVS3 encoding under specific scenarios, maximize the subjective quality of user-focused areas under limited bandwidth constraints, and expand the adaptability of the AVS3 standard in personalized encoding scenarios. Summary of the Invention
[0006] This invention addresses the shortcomings of existing AVS3 video coding technologies, such as lack of user intent guidance in bitrate allocation, insufficient accuracy in video visual region mapping, low flexibility in intra-frame and inter-frame bitrate scheduling, and an imbalance between coding quality and bitrate utilization. It proposes a differentiated AVS video coding method based on a user attention mechanism, which is a differentiated bitrate adaptation method based on an attention mechanism and a Grounding-DINO model.
[0007] This method takes user text intent as its core guide, introduces the Grounding-DINO model to achieve precise localization and matching of text intent with visual regions of video frames, constructs a dynamic bitrate scheduling system with attention region awareness, realizes differentiated bitrate allocation and reflow reuse of AVS3 video frame sequences in intra-frame spatial and inter-frame temporal dimensions, and ensures the matching degree between encoding quality and user intent through iterative verification and strategy correction mechanisms. While effectively improving bitrate resource utilization, it accurately guarantees the encoding quality of high attention regions, perfectly adapting to the personalized and intelligent AVS3 video encoding application needs.
[0008] The technical foundation of this invention relies on the AVS3 video coding standard, integrating core technologies such as semantic parsing of natural language processing, visual localization of open set object detection, inter-frame spatiotemporal feature correlation of computer vision, and construction of a preset redundant bitrate temporary buffer space for dynamic bitrate scheduling. The core innovation lies in using the Grounding-DINO model to achieve accurate mapping of user textual intent to video frame spatial regions, transforming abstract textual intent into a quantifiable and accurately localizable video frame attention region mask sequence and attention weights. Based on this, a dynamic bitrate scheduling system with intra-frame and inter-frame linkage is built, breaking through the technical limitations of fixed bitrate allocation, indiscriminate coding, and low accuracy of text-visual mapping in traditional AVS3 coding.
[0009] This invention achieves bitrate recirculation and intelligent allocation by constructing a preset redundant bitrate temporary buffer space, allowing the bitrate saved by high compression in low-attention areas to be accurately tilted to high-attention areas. At the same time, it introduces an iterative verification and strategy correction mechanism, integrating objective coding quality indicators and user subjective quality feedback to achieve dynamic optimization of the coding strategy, so that the final coding result is highly matched with the user's text intent, taking into account coding efficiency, bitrate utilization and personalized coding needs.
[0010] The method described in this invention takes user text intent as the guide, accurate visual positioning as the core, video frame attention region mask sequence as the scheduling basis, preset redundant bitrate temporary buffer space as the scheduling carrier, and iterative verification as the quality assurance, forming a complete AVS3 video differential bitrate adaptation system.
[0011] AVS is the general term for the entire set of standards developed by the my country Digital Audio and Video Coding Technology Standards Working Group, including AVS1, AVS2, and AVS3. The method of this invention is particularly suitable for AVS3.
[0012] An AVS3 video differential coding method based on user attention mechanism includes the following steps: Step 1: Initialize AVS3 encoding parameters and set the upper limit of the total encoding bitrate and the lower limit of the basic bitrate; Step 2: Obtain the user's input text intent and the AVS3 video frame sequence to be encoded. The AVS3 video frame sequence to be encoded includes I / P / B frames. Use the object detection model to generate a video frame attention region mask sequence with attention weights that is associated with the text intent. Step 3: Using the upper limit of the total encoding bitrate and the lower limit of the basic bitrate set in Step 1 as hard constraints, obtain the dynamic bitrate allocation plan in the intra-frame spatial dimension through the video frame attention region mask sequence, obtain the dynamic bitrate allocation plan in the inter-frame temporal dimension through the video frame attention region mask sequence and I / P / B frames, and generate an intra-frame-inter-frame linked bitrate scheduling index table through the dynamic bitrate allocation plan in the intra-frame spatial dimension and the dynamic bitrate allocation plan in the inter-frame temporal dimension. Step 4: Based on the bitrate scheduling index table of intra-frame and inter-frame linkage, perform differentiated coding on video frames with different attention weights in the video frame attention region mask sequence and the corresponding intra-frame coding units, and output the AVS3 initial encoded bitstream after completing the coding. Step 5: Perform attention region matching verification on video frames with different attention weights and different attention regions within the frames in the initial AVS3 encoded bitstream. Combine preset video quality indicators with user quality feedback for verification. Adjust the encoding strategy according to the verification results. Use the adjusted encoding strategy to correct the AVS3 video frame sequence to be encoded. Output the final AVS3 encoded bitstream to complete the enhanced display.
[0013] In step 1, the AVS3 encoding-related parameters include the frame group structure, coding unit division rules, and quantization parameter range specified by the AVS3 encoding standard.
[0014] In step 2, the target detection model adopts the Grounding-DINO open set target detection model.
[0015] In step 2, an object detection model is used to generate a sequence of attention-weighted video frame attention regions associated with the textual intent. Specifically, this includes: The object detection model extracts keywords of interest from textual intent, identifies corresponding target regions in the AVS3 video frame sequence to be encoded based on the keywords of interest, generates masks with attention weights, and obtains a video frame attention region mask sequence.
[0016] In step 4, the differential coding adopts a high-low compression ratio coding strategy. One strategy is to increase the quantization parameter and reduce the bit rate quota, and the other strategy is to reduce the quantization parameter and increase the bit rate quota.
[0017] In step 5, the video quality metrics include peak signal-to-noise ratio (PSNR) and structural similarity (SSIM).
[0018] In step 5, the verification method is to compare with the comprehensive matching degree threshold, which is determined based on video quality indicators and user quality feedback.
[0019] Specifically, the following steps are included: Step 1: Encoding Initialization. Following preset requirements, initialize the basic parameters related to the entire AVS3 encoding process, clearly setting the upper limit of the total encoding bitrate and the lower limit of the basic bitrate, and establishing a global bitrate control benchmark for differentiated encoding. Afterwards, perform preprocessing operations on the AVS3 video frame sequence to be encoded. Specifically, this includes: frame sequence denoising to eliminate interference information in the video frame sequence and prevent noise from affecting target positioning accuracy; inter-frame synchronization to align the temporal dimension of video frames; and resolution normalization to unify video frames of different resolutions to the preset encoding resolution, avoiding positioning and bitrate allocation deviations caused by resolution differences. Through the above preprocessing, ensure that the video frame quality meets the encoding requirements, providing high-quality input for subsequent steps. The preprocessing operations in this step are a prerequisite for ensuring the accuracy of subsequent semantic parsing, visual positioning, and inter-frame feature association.
[0020] Step 2: Intent parsing and attention region mask sequence generation. This step validates the user's input text intent information, focusing on the semantic integrity of the text and its relevance to the video frame sequence, eliminating invalid intent input. Validation effectively filters meaningless and irrelevant text input, ensuring the relevance and effectiveness of subsequent semantic parsing and visual localization.
[0021] After successful verification, the text intent information is semantically parsed based on a pre-trained natural language processing model to extract core semantic features and generate standardized text prompts. Next, the AVS3 video frame sequence to be encoded (including I-frames, P-frames, and B-frames) is acquired. The standardized text prompts and the pre-processed video frame sequence are input into the Grounding-DINO model. This model is an open-set object detection model, possessing the characteristic of achieving accurate matching between text and visual targets without fine-tuning for specific targets, and can adapt to diverse and personalized user text intent expressions. This model achieves accurate visual localization and bounding selection of target regions in video frames based on text intent. Simultaneously, optical flow combined with a convolutional neural network is used to extract the spatiotemporal features of the video frame sequence. By fusing the visual localization results and spatiotemporal features, a dual matching mapping between core semantic features and the spatial-temporal dimensions of video frames is achieved, generating an attention region mask sequence and defining the attention weights corresponding to each frame and region. The attention weights are divided into at least three levels: high, medium, and low. High attention corresponds to the core text scene and target region; medium attention corresponds to related scenes and regions, with weight values between high and low attention, serving as a balance between encoding quality and bitrate utilization; and low attention corresponds to irrelevant frames and background regions. This mask sequence clearly defines the spatial boundaries and weight values, providing high-precision basic input and feature support for the bitrate buffer space construction, bitrate scheduling, and subsequent differentiated encoding in step 3.
[0022] Step 3: Generate an intra-frame / inter-frame linked bitrate scheduling index table. Using the upper limit of the total coding bitrate and the lower limit of the basic bitrate set in Step 1 as hard constraints, construct an AVS3 preset redundant bitrate temporary storage buffer space. This buffer space is a scalable bitrate storage and scheduling carrier dynamically constructed based on bitrate allocation thresholds. The bitrate storage and scheduling thresholds can be adjusted in real time according to the length of the video frame sequence, the number of coding units, and the distribution of attention regions. Based on the precise attention region mask sequence and corresponding attention weights generated in Step 2, dynamic bitrate allocation planning is performed in two dimensions: in the intra-frame spatial dimension, dynamic bitrate scheduling is completed for CU (coding unit), TU (transform unit), and PU (prediction unit) under the AVS3 coding standard; in the inter-frame temporal dimension, dynamic bitrate scheduling is completed for AVS3 encoded I-frames, P-frames, and B-frames. When the coding bitrate consumption value of a low-attention region or low-attention frame is lower than the preset basic threshold, a bitrate reflow reuse mechanism is triggered, and the saved bitrate is returned to this buffer space in real time. Based on the above plan, a rate scheduling index table that links intra-frame and inter-frame operations is finally generated.
[0023] This index table clearly defines the upper and lower limits of bitrate allocation, reflow trigger conditions, and reuse priority for each coding unit, each type of video frame, and each precisely located target area. It is the core basis for achieving differentiated bitrate allocation.
[0024] Step 4: Differentiated Coding Implementation and Initial Bitstream Output. Based on the bitrate scheduling index table generated in Step 3, differentiated coding is implemented for frames, coding units, and precisely located target regions at different attention levels in the AVS3 video frame sequence. The coding units are the smallest coding units and merged coding units conforming to the AVS3 coding standard. The differentiated coding strategy strictly follows the technical specifications of the AVS3 coding standard, making targeted adjustments only to the prediction mode, quantization parameters, and bitrate allocation. A high compression ratio and low bitrate allocation coding strategy is adopted for low-attention frames and intra-frame low-attention regions, adapting to simplified intra-frame prediction modes and quantization parameters. The bitrate saved in this process is returned in real-time to the AVS3 preset redundant bitrate temporary storage buffer space. A low compression ratio and high bitrate allocation coding strategy is adopted for high-attention regions of high-attention frames, adapting to refined intra-frame prediction modes and quantization parameters. The bitrate in the buffer space is dynamically allocated to this region first, realizing the transfer and reuse of bitrate resources from low-value regions to high-value regions. After completing the differentiated coding processing of the entire frame sequence, the initial AVS3 encoded bitstream is output.
[0025] Step 5: Attention Region Matching Verification and Closed-Loop Correction. Attention region matching is verified for high-attention frames and core target regions in the initial AVS3 encoded bitstream output from Step 4. For verification metrics, Peak Signal-to-Noise Ratio (PSNR) and Structural Similarity (SSIM) are selected as objective quality indicators, combined with a subjective quality indicator based primarily on the Mean Opinion Score (MOS) of user ratings, to set a comprehensive matching threshold. The actual verification results are compared with the preset comprehensive threshold: if the matching degree does not reach the preset threshold, an iterative correction mechanism is initiated. This mechanism is a closed-loop optimization mechanism, which adjusts the attention weights of each region in the attention region mask sequence based on the matching degree deviation value, and optimizes the spatial boundaries of the mask regions in conjunction with the initial visual localization results. Subsequently, based on the updated mask sequence and weights, the intra-frame and inter-frame bitrate allocation coefficients and reflow multiplexing rules are recalculated, a full update of the bitrate scheduling index table is completed, and then the differential encoding process is re-executed from Step 3. This process ensures that each iteration brings the encoding result closer to the user's textual intent, effectively reducing the number of invalid iterations and improving overall encoding efficiency. If the verification is successful, the final AVS3 encoded bitstream will be output directly to complete the enhanced display.
[0026] In Step 1, the preprocessing operations for encoding initialization are a prerequisite for ensuring the accuracy of subsequent semantic parsing, visual localization, and inter-frame feature association. Denoising eliminates interference information in the video frame sequence, preventing noise from affecting target localization accuracy; frame synchronization aligns video frames along the temporal dimension; resolution normalization unifies video frames of different resolutions to a preset encoding resolution, avoiding localization and bitrate allocation deviations caused by resolution differences. Simultaneously, validating the user's text intent effectively filters out meaningless and irrelevant text input, ensuring the relevance and effectiveness of subsequent semantic parsing and visual localization.
[0027] Step 2: The Grounding-DINO model introduced in this step is an open-set object detection model. It possesses the characteristic of achieving accurate matching between text and visual targets without requiring fine-tuning for specific targets. It can adapt to diverse and personalized user text intent expressions, solving problems such as ambiguous visual region localization, missed labeling, and mislabeling in traditional semantic parsing and frame feature association. This results in clearer spatial boundaries and more accurate region segmentation of the attention region mask sequence. The attention weights are divided into at least three levels: high, medium, and low. The value of the medium attention weight falls between the high and low attention weight values. Its matched text intent associates scene frames and associated target regions, representing a balance between encoding quality and bitrate utilization. This ensures basic encoding quality without consuming excessive bitrate resources, achieving a layered and refined allocation of bitrate resources.
[0028] In step 3: The AVS3 preset redundant bitrate buffer space constructed in this step has dynamic expansion characteristics. It can adjust the bitrate storage and scheduling thresholds in real time according to the length of the video frame sequence, the number of coding units, and the distribution of attention regions, adapting to the coding needs of AVS3 video frame sequences of different sizes and types. The intra-frame and inter-frame linked bitrate scheduling index table is the core basis for realizing differentiated bitrate allocation. It clarifies the upper and lower limits of bitrate allocation, reflow trigger conditions, and reuse priorities for each coding unit, each type of video frame, and each precisely located target region, making bitrate scheduling more targeted and operable.
[0029] Step 4: The coding units involved in this step are the smallest coding units and merged coding units conforming to the AVS3 coding standard. The differentiated coding strategy strictly follows the technical specifications of the AVS3 coding standard, making targeted adjustments only to the prediction mode, quantization parameters, and bitrate allocation for different attention regions, ensuring the compatibility and universality of the coding results. The real-time bitrate reflow reuse allows the bitrate saved in low-attention regions to be immediately utilized by high-attention regions, maximizing the utilization efficiency of bitrate resources and avoiding idle and wasted bitrate resources.
[0030] In step 5: The iterative correction mechanism in this step is a closed-loop optimization mechanism. Its core is based on the matching degree deviation value, combined with the initial visual localization results, to accurately update the attention region mask sequence and the bitrate scheduling index table, rather than making random adjustments. This ensures that each iteration brings the encoding result closer to the user's text intent, effectively reducing the number of invalid iterations and improving overall encoding efficiency. Furthermore, by combining objective and subjective quality indicators to set a comprehensive matching degree threshold, the objectivity of the encoding quality is guaranteed while also taking into account the user's personalized subjective needs. This ensures that the final AVS3 encoded bitstream both meets technical standards and accurately matches the user's intent.
[0031] This invention's method is based entirely on the AVS3 video coding standard. All coding strategies and scheduling mechanisms are deeply integrated with AVS3 coding technology, ensuring the standardization and compatibility of the coding results. Simultaneously, guided by user textual intent, it achieves precise text-visual mapping through the Grounding-DINO model and intelligent, differentiated bitrate allocation through an attention mechanism. This addresses the industry pain points of traditional AVS3 coding, such as insufficient coding quality in highly important areas due to indiscriminate coding, low accuracy of text-visual mapping, and low bitrate utilization. This invention reduces overall bitrate consumption and improves bitrate utilization while precisely ensuring the coding quality of user-focused areas, making AVS3 video coding more suitable for personalized and intelligent application needs. It can be widely applied to various video coding scenarios, including short video coding, high-definition live video streaming, video-on-demand, and intelligent surveillance video coding.
[0032] Compared with the prior art, the present invention has the following advantages: This invention proposes a differentiated bitrate adaptation method based on an attention mechanism and the Grounding-DINO model, which deeply aligns with the AVS3 video coding standard. It specifically addresses the core pain points of traditional coding techniques, such as lack of user intent guidance in bitrate allocation, insufficient accuracy in text-to-visual region mapping, rigid intra-frame and inter-frame scheduling, and an imbalance between coding quality and bitrate utilization. This method centers on user text intent, using the Grounding-DINO model to achieve precise anchoring between text and video visual regions. It constructs a hierarchical bitrate allocation system using an attention mechanism, achieves dynamic bitrate reflow and reuse through a pre-set redundant bitrate buffer, and optimizes coding performance with a closed-loop iterative verification mechanism. It is fully compatible with the AVS3 coding specification, requires no large-scale reconstruction of the existing coding framework, has low deployment costs, and is applicable to a wide range of scenarios. It ensures the coding quality of user-focused core regions while minimizing bitrate consumption in non-critical regions, significantly improving bitrate resource utilization. It provides an efficient solution for personalized video coding scenarios such as short videos, high-definition live streaming, and intelligent monitoring, possessing strong engineering application value and technological innovation.
[0033] The innovativeness of this invention is specifically reflected in the following aspects: 1) This invention innovatively proposes an attention mask construction scheme based on "text intent-driven + precise visual positioning," breaking through the limitations of traditional AVS3 encoding's indiscriminate bitrate allocation. In existing technologies, bitrate allocation is mostly based on fixed parameters such as frame type and resolution, which are disconnected from the actual user attention needs. Furthermore, the mapping between text intent and video visual regions is ambiguous, easily leading to insufficient bitrate in core areas and wasted bitrate in redundant areas. This invention introduces the Grounding-DINO model to achieve precise matching between text and video targets, transforming abstract text intent into a quantifiable sequence of attention region masks. Combined with inter-frame spatiotemporal feature correlation, it achieves spatial-temporal dual-dimensional attention layering, allowing bitrate allocation to precisely match user needs and solving the industry pain point of "mismatch between bitrate investment and user attention."
[0034] 2) This invention innovatively constructs a pre-defined redundant bitrate buffer space and a reflow reuse mechanism that links intra-frame and inter-frame operations, improving bitrate resource utilization. Traditional AVS3 encoding uses a fixed bitrate budget allocation mode, where the bitrate quotas for each region within a frame and for different types of frames (I / P / B frames) between frames are relatively fixed. Bitrate consumed in low-importance regions cannot be reused, resulting in low overall bitrate efficiency. This invention constructs a dynamically expandable pre-defined redundant bitrate buffer space. Based on attention weights, it implements differentiated scheduling for intra-frame CU / TU / PU units and different types of frames between frames. Bitrate saved by high compression in low-attention regions is returned to the pre-defined redundant bitrate buffer space in real time and prioritized for allocation to core regions, achieving "on-demand allocation and cyclic reuse" of bitrate. Under the same bitrate budget, this can significantly improve the encoding quality of core regions or reduce overall bitrate consumption while ensuring core quality.
[0035] 3) This invention innovatively designs a closed-loop iterative verification mechanism of "objective indicators + subjective feedback" to achieve dynamic optimization of the encoding strategy. Existing AVS3 encoding quality assessments mostly rely on single objective indicators (such as PSNR, SSIM), ignoring users' subjective visual needs, and lack a strategy correction stage after encoding, making it difficult to adapt to personalized needs in diverse scenarios. This invention combines objective quality indicators with user subjective ratings (MOS) to set a comprehensive threshold, performs matching degree verification on high-attention core areas, and iteratively optimizes attention masks and bitrate scheduling strategies based on deviation values when the threshold is not met, forming a closed-loop system of "encoding-verification-correction". This ensures that the final encoding result not only meets technical standards but also accurately matches the user's subjective experience, filling the technical gap of "disconnect between technical indicators and subjective needs" in personalized encoding scenarios. Attached Figure Description
[0036] Figure 1 This is a schematic diagram of the overall architecture of the method.
[0037] Figure 2 A schematic diagram of the preset redundant bitrate temporary buffer space and scheduling process for AVS3. Detailed Implementation
[0038] The attention-based differential bitrate adaptation method of the present invention will be further described below with reference to the accompanying drawings, in which: Figure 1 This is a schematic diagram of the overall architecture of the method. Figure 2 A schematic diagram of the preset redundant bitrate temporary buffer space and scheduling process for AVS3.
[0039] Combination Figure 1 The attention-based differential bitrate adaptation method has an overall architecture consisting of four layers from bottom to top: a data preprocessing layer, a core computation layer, a bitrate scheduling layer, and a result verification layer. These layers work together to achieve a closed loop from input to final encoded bitstream output. 1) The data preprocessing layer is responsible for AVS3 video frame sequence denoising, frame synchronization, resolution normalization, and user text intent validity verification, providing standardized input for subsequent calculations; 2) The core computing layer includes modules for text semantic parsing, precise visual positioning, inter-frame spatiotemporal feature association, and attention mask generation, which are the core for realizing the mapping between text intent and video regions; 3) The bitrate scheduling layer consists of a preset redundant bitrate temporary storage buffer space, intra-frame and inter-frame bitrate scheduling, and a differentiated coding module, and is responsible for the dynamic allocation and reflow reuse of bitrate. 4) The result verification layer enables objective evaluation of coding quality, collection of user subjective feedback, and iterative correction of strategies to ensure that the coding results match the user's intent.
[0040] The process for generating the attention region mask sequence is as follows: 1) The module input is a preprocessed AVS3 video frame sequence and standardized text prompts, and the output is a sequence of attention region masks with weighted annotations; 2) First, the text intent is semantically parsed using a pre-trained BERT model to extract core semantic features and generate standardized prompt words. These are then input into the Grounding-DINO model to accurately locate the target region in the text and video frames, and output the target region's coordinate box. 3) Optical flow is used to extract the inter-frame spatiotemporal features of the video. Formula (1) is the core formula for calculating the optical flow field, where At time t coordinate pixel value, , These are the optical flow vectors in the x and y directions, respectively. For pixel gradient, Inter-frame temporal gradient: 4) The Grounding-DINO localization results are integrated with the inter-frame spatiotemporal features. Feature mapping is completed through a convolutional neural network to generate an initial mask sequence. Then, high, medium and low attention weights are assigned based on semantic relevance. Finally, an accurate mask sequence is output. The default weight values are high = 0.9-1.0, medium = 0.3-0.8 and low = 0.1-0.3, which can be dynamically adjusted according to the scene.
[0041] The AVS3 preset redundant code rate temporary storage buffer space and scheduling process are as follows: 1) The preset redundant bitrate temporary buffer space is a dynamically expandable structure. The initial capacity is set based on the total video bitrate budget, and the expansion threshold is 80% of the current capacity. When the bitrate usage reaches the threshold, the capacity will automatically expand by 2 times. 2) The scheduling process is divided into intra-frame spatial scheduling and inter-frame temporal scheduling: For CU / TU / PU units encoded in AVS3 within a frame, the bit rate is allocated according to the attention weight. CU units in high-attention regions are given priority to be allocated high bit rates, and 4×4 / 8×8 fine-grained partitioning is adopted; For I-frames, P-frames, and B-frames between frames, I-frame coding resources are allocated first for high-attention frames, and B-frame simplified coding is adopted for low-attention frames. 3) The trigger condition for bitrate reflow reuse is that the bitrate consumption of low attention region / frame is lower than the preset base threshold (which is 30% of the single frame bitrate budget by default). The bitrate real-time reflow is saved into the preset redundant bitrate temporary storage buffer space and allocated to high attention region according to attention priority. 4) The calculation of the bitrate allocation coefficient is shown in formula (2), where This is the attention weight coefficient. Frame type weights (I-frame = 1.2, P-frame = 0.9, B-frame = 0.7). To allocate bitrate, Budget for total bitrate per frame: As shown in Table 1, the encoding strategies corresponding to the bitrate scheduling index table are as follows: 1) The encoding strategy is based on the rate scheduling index table, and the encoding parameters are configured differently according to the attention level, strictly following the AVS3 encoding standard; 2) High attention region coding configuration: adopts low compression ratio (quantization parameter QP=22-28), intra-frame prediction mode is AMP (adaptive multi-partition prediction), and SAO (sample adaptive offset) optimization is enabled to ensure detail quality; 3) Medium attention region coding configuration: adopt medium compression ratio (QP=29-35), intra-frame prediction mode is IPCM (intra-frame prediction coding mode), and some redundant optimization modules are turned off to balance quality and efficiency; 4) Low attention region coding configuration: High compression ratio (QP=36-42), intra-frame prediction mode to simplify DCT transform, and fast skip mode enabled to maximize bitrate savings; 5) During the encoding process, the bitrate consumption of each region is statistically analyzed in real time, and the saved bitrate is returned to the preset redundant bitrate temporary storage buffer space and the scheduling index table is updated synchronously.
[0042] The iterative verification and policy correction mechanism is as follows: 1) The verification objects are high-attention frames and core target regions in the initial encoded bitstream. The verification indicators are divided into objective indicators and subjective indicators. 2) The objective indicators are PSNR, SSIM, and MS-SSIM, and the calculation methods are shown in formulas (3) and (4) respectively. , These are the pixel matrices of the original frame and the encoded frame, respectively. M and N are the height and width of the image (in pixels), respectively, and 255 is the maximum gray level of the 8-bit pixel value. The average pixel value. For covariance, , For variance, , For constants: 3) Subjective metrics use the Mean Opinion Score (MOS) (1-5 points) to collect user ratings of the visual quality of the core area; 4) The overall matching degree threshold is set as the sum of the objective indicator weighted value (accounting for 70%) and the MOS value (accounting for 30%). If the threshold is not reached, the attention weight and mask boundary are adjusted based on the deviation value gradient, the bitrate scheduling index table is updated, and the encoding stage is backtracked to re-execute.
[0043] The initial settings for the operation parameters of this method are as follows: 1) The operation types cover five major categories: semantic parsing, visual localization, optical flow calculation, bit rate scheduling, and differential coding; 2) The core parameters are set as follows: BERT model hidden layer dimension = 768, Grounding-DINO confidence threshold = 0.7, optical flow calculation window size = 15×15, initial capacity of preset redundant bit rate temporary buffer space = 5Mbps, QP initial range = 22-42, maximum number of iterations and verifications = 3 times. 3) The output results include attention mask sequence (size consistent with video frame), bitrate scheduling index table (including unit-level bitrate allocation rules), and initial / final AVS3 bitstream (supports .ts / .mp4 format encapsulation).
[0044] Traditional coding suffers from low intent matching scores due to the lack of user intent guidance, while this invention achieves a high degree of consistency between the coding results and user needs through precise visual positioning and iterative verification.
[0045] The above are merely specific embodiments of the present invention and are not intended to limit the invention. Those skilled in the art should recognize that any adjustments or modifications to the operational parameters, module structure optimizations, or encoding strategy fine-tuning of the present invention will fall within the protection scope of the present invention.
[0046] Table 1 Rate Scheduling Index Table
Claims
1. An AVS video differential coding method based on user attention mechanism, characterized in that, Includes the following steps: Step 1: Initialize AVS3 encoding parameters and set the upper limit of the total encoding bitrate and the lower limit of the basic bitrate; Step 2: Obtain the user's input text intent and the AVS3 video frame sequence to be encoded. The AVS3 video frame sequence to be encoded includes I / P / B frames. Use the object detection model to generate a video frame attention region mask sequence with attention weights that is associated with the text intent. Step 3: Using the upper limit of the total encoding bitrate and the lower limit of the basic bitrate set in Step 1 as hard constraints, obtain the dynamic bitrate allocation plan in the intra-frame spatial dimension through the video frame attention region mask sequence, obtain the dynamic bitrate allocation plan in the inter-frame temporal dimension through the video frame attention region mask sequence and I / P / B frames, and generate an intra-frame-inter-frame linked bitrate scheduling index table through the dynamic bitrate allocation plan in the intra-frame spatial dimension and the dynamic bitrate allocation plan in the inter-frame temporal dimension. Step 4: Based on the bitrate scheduling index table of intra-frame and inter-frame linkage, perform differentiated coding on video frames with different attention weights in the video frame attention region mask sequence and the corresponding intra-frame coding units, and output the AVS3 initial encoded bitstream after completing the coding. Step 5: Perform attention region matching verification on video frames with different attention weights and different attention regions within the frames in the initial AVS3 encoded bitstream. Combine preset video quality indicators with user quality feedback for verification. Adjust the encoding strategy according to the verification results. Use the adjusted encoding strategy to correct the AVS3 video frame sequence to be encoded. Output the final AVS3 encoded bitstream to complete the enhanced display.
2. The AVS video differential coding method based on user attention mechanism according to claim 1, characterized in that, In step 1, the AVS3 encoding-related parameters include the frame group structure, coding unit division rules, and quantization parameter range specified by the AVS3 encoding standard.
3. The AVS video differential coding method based on user attention mechanism according to claim 1, characterized in that, In step 2, the target detection model adopts the Grounding-DINO open set target detection model.
4. The AVS video differential coding method based on user attention mechanism according to claim 1, characterized in that, In step 2, an object detection model is used to generate a sequence of attention-weighted video frame attention regions associated with the textual intent. Specifically, this includes: The object detection model extracts keywords of interest from textual intent, identifies corresponding target regions in the AVS3 video frame sequence to be encoded based on the keywords of interest, generates masks with attention weights, and obtains a video frame attention region mask sequence.
5. The AVS video differential coding method based on user attention mechanism according to claim 1, characterized in that, In step 4, the differential coding adopts a high-low compression ratio coding strategy. One strategy is to increase the quantization parameter and reduce the bit rate quota, and the other strategy is to reduce the quantization parameter and increase the bit rate quota.
6. The AVS video differential coding method based on user attention mechanism according to claim 1, characterized in that, In step 5, the video quality metrics include peak signal-to-noise ratio (PSNR) and structural similarity (SSIM).
7. The AVS video differential coding method based on user attention mechanism according to claim 1, characterized in that, In step 5, the verification method is to compare with the comprehensive matching degree threshold, which is determined based on video quality indicators and user quality feedback.