An efficient image and video coding rate control method based on deep learning

By constructing a dual-branch heterogeneous prediction structure and dynamic correction using Lagrange multipliers, combined with decoupled perception of spatiotemporal features and hard-gated switching for sudden motion, the prediction divergence problem of existing video coding rate control methods in complex scenarios is solved. This achieves accurate allocation of bitrate resources and stable control of buffer levels, improving the adaptability and stability of video coding.

CN122372736APending Publication Date: 2026-07-10JIANGSU YUNBO INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
JIANGSU YUNBO INFORMATION TECH CO LTD
Filing Date
2026-04-29
Publication Date
2026-07-10

AI Technical Summary

Technical Problem

Existing deep learning-based video coding rate control methods struggle to achieve comprehensive predictive modeling when faced with challenges such as frequent motion mutations, drastic bandwidth fluctuations, and significant differences in the human eye's visual sensitivity areas. This results in insufficient adaptive optimization capabilities for rate allocation strategies and a lack of clear intervention at the encoder's underlying level, impacting the practical value and stability of video coding.

Method used

A dual-branch heterogeneous prediction structure is constructed, and a dynamic correction mechanism using Lagrange multipliers and a virtual buffer verification gating mechanism are introduced. Combined with spatiotemporal feature decoupling perception and a motion burst hard gating switching mechanism, the correlation between local texture and global bitrate budget is mined through a nonlinear mapping network to achieve accurate allocation of bitrate resources and stable control of buffer level.

Benefits of technology

It significantly improves the accuracy of error compensation and the robustness of the system in preventing out-of-bounds errors. It can achieve precise allocation of bitrate resources and stable control of buffer level in complex video scenarios, and enhance the adaptability and stability of video encoding.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122372736A_ABST
    Figure CN122372736A_ABST
Patent Text Reader

Abstract

The application discloses a kind of high-efficiency image video coding rate control methods based on deep learning, comprising: S1, current image frame and front state parameter are collected;S2, decoupling construction feature atlas is carried out to space-time characteristics and is fused, and output double-flow heterogeneous representation tensor;S3, by nonlinear mapping, lagrange multiplier dynamic adjustment coefficient is output;S4, the coefficient is used to correct reference multiplier to solve output target quantization parameter;S5, according to actual encoding error and buffer fullness, input improved NMPC algorithm, introduce hard-gate switching mechanism to deduce error trajectory, output closed-loop residual compensation bias;S6, bias is superimposed to generate modified bit budget, if buffer fullness is out of bounds, then trigger adaptive redistribution algorithm to carry out secondary correction to complete closed-loop control.The application effectively improves the perceptual rationality and system robustness of code rate allocation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of cross-integration of video coding and deep learning, and in particular to an efficient image and video coding rate control method based on deep learning. Background Technology

[0002] Deep learning-based rate control methods have been widely applied in recent years to ultra-high-definition live streaming, cloud gaming streaming, and intelligent security monitoring due to their perceptual modeling and dynamic resource allocation capabilities during video compression. They have become an important development direction for improving video transmission quality and system stability. However, in practical applications, video encoding scenarios face numerous challenges, such as frequent motion abrupt changes, drastic bandwidth fluctuations, and significant differences in the human eye's visual sensitivity areas. Therefore, the deployment effectiveness of deep learning-based rate control methods remains constrained by various factors.

[0003] Most current video rate control methods rely on a single-dimensional state input, making it difficult to fully utilize multi-source information such as spatiotemporal texture features, buffer level changes, and historical coding errors. This results in a lack of comprehensiveness in predicting and modeling bitrate fluctuations. Some systems only use a fixed rate-distortion model to calculate quantization parameters, ignoring the combined influence of multiple factors such as local visual sensitivity, global bitrate budget correlation, and coding tree unit mode selection costs, thus limiting the adaptive optimization capability of bitrate allocation strategies. Furthermore, the mapping process from deep features to quantization parameters lacks a transparent interpretive path, making it difficult to provide clear intervention criteria to the encoder layer or system control end, affecting the reliability and controllability of rate control results.

[0004] Furthermore, most existing rate control prediction mechanisms in error compensation and buffer adjustment are static and single designs, failing to dynamically switch prediction branches and solution strategies according to sudden changes in scene motion. This results in severe divergence of the prediction trajectory in scenes with rapid shot changes or violent motion, making it difficult to adapt to the continuous changes and evolution of video content. In addition, the lack of a hard constraint safety fallback mechanism to deal with extreme budget deviations seriously affects the practical value of the encoder and its stability in preventing out-of-bounds errors in real and complex scenarios.

[0005] Therefore, how to provide an efficient image and video coding rate control method based on deep learning is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention

[0006] One objective of this invention is to propose a high-efficiency image and video coding rate control method based on deep learning. This invention fully integrates key steps such as spatiotemporal feature decoupling perception, dynamic correction using Lagrange multipliers, error trajectory deduction, and virtual buffer verification gating. It constructs an intelligent rate control process with local texture association, low-level coding intervention, advance prediction of temporal errors, and mandatory fallback for budget overruns, achieving precise allocation of bitrate resources and stable control of buffer levels in complex video scenarios. The core improvement of this invention lies in constructing a dual-branch heterogeneous prediction structure and introducing a hard-gating switching mechanism based on motion abruptness, embedding it into an improved NMPC algorithm. For scenes with sudden motion changes, it decisively cuts off stable channels and activates the abrupt branch, completely solving the defect of prediction divergence in traditional single models when the image changes drastically. Combined with an adaptive bit-forced reallocation algorithm based on visual perception weights, it significantly improves the accuracy of error compensation and the system's robustness against overruns, thus effectively solving problems such as one-sided state modeling, lagging response to sudden changes, and easy imbalance in budget control in existing methods.

[0007] An efficient image and video coding rate control method based on deep learning according to an embodiment of the present invention includes the following steps: S1. Acquire the current image frame to be encoded and the previous state parameters in the video encoding pipeline, and output the current image frame and the previous state parameters. S2. Perform downsampling and spatiotemporal feature decoupling processing on the current image frame to construct a spatial perception feature map reflecting human visual sensitivity and a temporal motion complexity feature map reflecting coding difficulty, respectively. Perform cross-modal attention fusion and output a dual-stream heterogeneous representation tensor. S3. The dual-stream heterogeneous representation tensor is concatenated with the previous state parameters. The correlation between local texture and global bit rate budget in the current image frame is mined through nonlinear mapping, and the dynamic adjustment coefficient of the Lagrange multiplier is output. S4. The reference Lagrange multipliers in the pre-state parameters are weighted and corrected using the Lagrange multipliers dynamic adjustment coefficients to generate the target Lagrange multipliers, which are then injected into the rate-distortion optimization layer of the video encoder for secondary solution, thereby indirectly deriving and outputting the target quantization parameters of each coding tree unit. S5. Calculate the instantaneous coding error after actual coding based on the target quantization parameters, and input the improved NMPC algorithm with the initial buffer fullness to construct a dual-branch heterogeneous prediction structure and introduce a hard-gating switching mechanism based on motion bursts. Use the variance of the motion vector field distribution as the criterion to perform hard cut-off and activation-only operations on the two prediction branches, deduce the future error trajectory, generate candidate compensation sequences, and output the closed-loop residual compensation bias. S6. The closed-loop residual compensation bias is superimposed on the bit budget reference plane of the next image frame to be encoded to generate the corrected bit budget. If the initial buffer fullness is detected to exceed the safe threshold range, the adaptive bit forced reallocation algorithm is triggered to perform secondary correction and complete efficient closed-loop rate control.

[0008] Optionally, S1 specifically includes: S11. Extract the pixel displacement data generated in the forward motion estimation process in the video coding pipeline, arrange them into a two-dimensional matrix according to the original image resolution, and generate and output the motion vector field distribution. S12. Read the value of the current remaining bits in the buffer and the total capacity of the buffer during the video encoding pipeline operation, divide the current remaining bits by the total capacity of the buffer to calculate the ratio, generate and output the initial buffer fullness; S13. Divide the preset target bit rate value by the video frame rate value to calculate the target bit value of a single frame. Obtain the difference between the original pixel value and the reconstructed pixel value of all pixel blocks in the current image frame under the initial quantization parameters. Square the difference and add them together to obtain the absolute error sum of squares value. Divide the target bit value of a single frame by the absolute error sum of squares value to obtain the slope value. Set the slope value as the reference Lagrange multiplier and output it. S14. Combine the pixel matrix, motion vector field distribution, initial buffer fill value, and reference Lagrange multiplier value of the current image frame to be encoded into the same data set in order, and output the current image frame and the previous state parameters.

[0009] Optionally, S2 specifically includes: S21. Input the pixel matrix of the current image frame into a downsampling convolutional layer of a preset size, add the gray values ​​of every two adjacent pixels in the horizontal direction and calculate the average value, add the gray values ​​of every two adjacent pixels in the vertical direction and calculate the average value, and generate a downsampling image reduced to the original preset ratio size. S22. Divide the downsampled image into multiple image blocks according to a preset pixel size. Subtract the average gray value of all pixels in each image block from the average gray value of the image block and square the result. Add the squared results of all pixels and divide by the total number of pixels to calculate the variance value. Arrange the variance values ​​of all image blocks into a two-dimensional matrix in row and column order to generate a spatial perception feature map. S23. Divide the current image frame into multiple image blocks according to the preset pixel size, extract the horizontal displacement value and vertical displacement value of the image block corresponding to the motion vector field distribution in the previous state parameters, add the horizontal and vertical displacement values ​​of all pixels in the same image block and take the absolute value, and add them to obtain the motion amplitude value. Arrange the motion amplitude values ​​of all image blocks into a two-dimensional matrix in the row and column order to generate a temporal motion complexity feature map. S24. Multiply the variance value of each image block in the spatial perception feature map with the motion amplitude value of the corresponding image block in the temporal motion complexity feature map to obtain the correlation value. Divide each variance value in the spatial perception feature map by the sum of all variance values ​​to obtain the first weight. Divide each motion amplitude value in the temporal motion complexity feature map by the sum of all motion amplitude values ​​to obtain the second weight. S25. Multiply the associated values ​​by the product of the first weight and the second weight, arrange all the multiplication results in their original positions and add them together to generate and output the dual-stream heterogeneous representation tensor.

[0010] Optionally, S3 specifically includes: S31. Expand the dual-stream heterogeneous representation tensor into a one-dimensional feature sequence along the channel dimension, extract the initial buffer fullness and the benchmark Lagrange multiplier from the pre-state parameters, and multiply them by the corresponding preset scaling factor respectively. Then, append the two scaled values ​​to the end of the one-dimensional feature sequence in sequence to generate a global state enhancement feature vector. S32. Input the global state enhancement feature vector into the first reshaping layer of the nonlinear mapping network, rearrange the global state enhancement feature vector into a three-dimensional tensor, divide the three-dimensional tensor into multiple local windows along the depth dimension, calculate the local bias by subtracting the average value of all values ​​in the window from the value at each position in each local window, and add the local bias to the original value to output the local texture self-calibration tensor. S33. Input the local texture self-calibration tensor into the cross-attention layer of the nonlinear mapping network, divide the local texture self-calibration tensor into a query matrix, a key matrix and a value matrix, transpose the query matrix and the key matrix, multiply them and divide by the preset square root value to obtain the similarity weight matrix, input the value of each row in the similarity weight matrix into the Softmax function to calculate the exponential ratio, multiply the exponential ratio with the value matrix and sum over the depth dimension, and output the local global association tensor. S34. Input the local-global correlation tensor into the two one-dimensional convolutional layers of the feedforward network layer of the nonlinear mapping network. Between the two one-dimensional convolutional layers, divide each value by a preset ratio and add the corresponding value to calculate the normalized value. After the second one-dimensional convolutional layer, perform Gaussian error linear unit activation operation on each value to map the negative value to an exponentially decaying output value and output the nonlinear mapping tensor. S35. Input the nonlinear mapping tensor into the output layer of the nonlinear mapping network, and process it sequentially through one-dimensional convolution operation and Sigmoid activation function to generate a continuous floating-point value sequence. Divide the current image frame into multiple grid regions according to the preset division size of the coding tree unit, and assign the floating-point values ​​in the continuous floating-point value sequence to the corresponding grid regions as Lagrange multiplier dynamic adjustment coefficients and output them.

[0011] Optionally, S4 specifically includes: S41. Extract the dynamic adjustment coefficients of the Lagrange multipliers corresponding to each coding tree unit in the current image frame, read the reference Lagrange multipliers corresponding to the same coding tree unit in the pre-state parameters, multiply the dynamic adjustment coefficients of the Lagrange multipliers directly with the reference Lagrange multipliers, generate and output the target Lagrange multipliers of the corresponding coding tree unit. S42. Input the target Lagrange multiplier of each coding tree unit into the rate-distortion optimization layer of the video encoder. In the rate-distortion optimization layer, obtain the sum of squared absolute errors between the original pixel value and the predicted pixel value of the corresponding coding tree unit in the current prediction mode. Multiply the sum of squared absolute errors by the target Lagrange multiplier to obtain the distortion cost value. S43. For the prediction mode index, motion vector data and residual matrix data generated by the corresponding coding tree unit in the current prediction mode, convert the prediction mode index, motion vector data and residual matrix data into binary bit streams in sequence, count the total number of bits contained in the binary bit stream as the actual bit value of encoding consumption, and directly add the distortion cost value to the actual bit value of encoding consumption to calculate the total distortion cost value. S44. Traverse all possible prediction modes of the corresponding coding tree unit and repeat the step of calculating the total rate distortion cost value a preset number of times. Compare all total rate distortion cost values, find the optimal prediction mode corresponding to the total rate distortion cost value with the smallest value, and extract the residual data under the optimal prediction mode. S45. Input the residual data in the optimal prediction mode into the quantizer, divide the value of the residual data by the initial quantization step size and round down to obtain the integer coefficient, multiply it by the initial quantization step size to obtain the quantized residual value, and calculate the difference between the value of the residual data and the quantized residual value as the quantization error value. S46. Multiply the quantization error value by the preset compensation coefficient to obtain the compensation offset value, subtract the compensation offset value from the initial quantization step size to obtain the target quantization step size, divide the target quantization step size by the preset logarithmic base to obtain the logarithmic value, and add the logarithmic value to the preset offset constant to obtain the target quantization parameter.

[0012] Optionally, S5 specifically includes: S51. Extract the single-frame target bit value allocated in the current image frame and the actual number of bits consumed after actual encoding. Subtract the actual number of bits consumed from the single-frame target bit value to calculate the difference and generate the instantaneous encoding error. S52. Extract the initial buffer fullness from the pre-state parameters, concatenate the instantaneous coding error value with the initial buffer fullness value in the order of front and back to generate a two-dimensional state vector, input the two-dimensional state vector into the improved NMPC algorithm, construct a dual-branch heterogeneous prediction structure inside the improved NMPC algorithm, set up the first calculation channel containing two one-dimensional convolutional layers as the stationary prediction branch, and set up the second calculation channel containing three fully connected layers with the number of output nodes decreasing sequentially as the mutation prediction branch. S53. Extract the motion vector field distribution from the previous state parameters. Subtract the corresponding average horizontal displacement and average vertical displacement from the horizontal displacement values ​​and vertical displacement values ​​of all pixels in the motion vector field distribution. Square the results of the subtraction and sum them up. Then divide by the total number of pixels in the motion vector field distribution to calculate the variance. S54. The calculated variance is compared with the preset mutation threshold as the motion burst criterion. A hard gating switching mechanism is executed. When the variance is greater than or equal to the preset mutation threshold, it is determined to be a motion burst. The values ​​of all positions in the two one-dimensional convolutional layers in the first calculation channel are directly replaced with 0 to perform a hard cut-off operation. The numerical calculation process of the three fully connected layers in the second calculation channel is retained, and only the mutation prediction branch is activated. S55. When the variance is less than the preset mutation threshold, the motion is judged to be stable. The values ​​of all positions in the three fully connected layers in the second calculation channel are directly replaced with 0 to perform a hard cut-off operation, while the numerical calculation process of the two one-dimensional convolutional layers in the first calculation channel is retained, and only the stable prediction branch is activated. S56. Generate future error trajectories in the only activated stationary prediction branch or abrupt prediction branch, and extract the number of compensation bits ranked first in the calculated candidate compensation sequence. Use the extracted number of compensation bits directly as the closed-loop residual compensation bias of the next frame and output it.

[0013] Optionally, S56 specifically includes: S561. In the stationary prediction branch that is activated only, the two-dimensional state vector is input into the first one-dimensional convolutional layer. The value of the one-dimensional convolutional kernel at each position is multiplied and added with the value at the corresponding position in the two-dimensional state vector. After passing through the ReLU activation function, the one-dimensional intermediate vector is output and input into the second one-dimensional convolutional layer to repeat the operation of multiplication, addition and replacement of negative numbers, and output the predicted value at the current time. S562. The current predicted value plus the instantaneous coding error is used as the two-dimensional state vector of the next moment and input into two one-dimensional convolutional layers. The summation input operation is performed continuously a preset number of times. The predicted values ​​output each time are arranged in chronological order to generate the future error trajectory. S563. In the mutation prediction branch that is only activated, the two-dimensional state vector is input into the first fully connected layer for fully connected layer linear mapping, the first bias value is added, the hyperbolic tangent function is applied, and the calculation result is compressed to between -1 and 1 to output the first hidden vector. After inputting into the second fully connected layer to output the second hidden vector, the second hidden vector is input into the third fully connected layer to output the predicted value at the current time. S564. The current predicted value plus the instantaneous coding error is used as the two-dimensional state vector of the next moment and input into three fully connected layers. The summation input operation is performed continuously a preset number of times. The predicted values ​​output each time are arranged in chronological order to generate the future error trajectory. S565. In the improved NMPC algorithm, the upper limit and lower limit of the buffer fullness corresponding to each future image frame are preset as the constraint boundary of the rolling optimization problem with hard constraints. For the preset number of predicted values ​​in the future error trajectory, the same preset number of compensation bits to be obtained are initialized with an initial value of 0. The future buffer prediction value is calculated by subtracting the corresponding compensation bit number from each predicted value. S566. Determine whether each future buffer prediction value is between the upper limit of buffer fullness and the lower limit of buffer fullness. When it exceeds the value range, calculate the difference between the future buffer prediction value and the upper limit of buffer fullness or the lower limit of buffer fullness as the out-of-bounds deviation value. Multiply the out-of-bounds deviation value by the preset penalty weight to obtain the penalty cost value. S567. The total cost value is calculated by adding the sum of the squares of the preset number of compensation bits to all penalty cost values. Within the preset search step size, the value of the preset number of compensation bits is changed one by one and the total cost value is calculated repeatedly. All calculation results are compared, and the 5 compensation bits that minimize the total cost value are extracted and arranged in chronological order to generate a candidate compensation sequence. S568. In the generated candidate compensation sequence, extract the number of compensation bits ranked first, and directly use the extracted number of compensation bits as the closed-loop residual compensation bias of the next frame and output it.

[0014] Optionally, S6 specifically includes: S61. Extract the single-frame target bit value corresponding to the current image frame as the bit budget reference surface for the next image frame to be encoded, read the closed-loop residual compensation bias, add the value of the bit budget reference surface to the value of the closed-loop residual compensation bias, and calculate and generate the corrected bit budget. S62. Extract the total number of coding tree units contained in the next image frame to be encoded, divide the corrected bit budget by the total number of coding tree units to calculate the initial allocated number of bits for each coding tree unit, subtract the corrected bit budget from the current remaining number of bits and add the actual number of bits consumed to calculate the predicted remaining number of bits, and divide the predicted remaining number of bits by the total capacity of the buffer to calculate the prediction buffer fullness. S63. The upper and lower limits of the preset safety threshold range are compared with the fullness of the prediction buffer and the safety threshold range. When the fullness of the prediction buffer is between the upper and lower limits, the corrected bit budget is directly used as the final bit budget output to complete efficient closed-loop rate control. When the fullness of the prediction buffer is greater than the upper limit or less than the lower limit, the virtual buffer verification gating mechanism is triggered. S64. When the virtual buffer verification gating mechanism is triggered, extract the spatial perception feature map corresponding to the current image frame, divide the variance value of each image block in the spatial perception feature map by the sum of the variance values ​​of all image blocks, and calculate and generate the visual perception weight corresponding to each coding tree unit. S65. Trigger the adaptive bit forced reallocation algorithm based on visual perception weights, multiply the visual perception weight of each coding tree unit by the corrected bit budget to calculate the weighted budget value, and add the weighted budget values ​​of all coding tree units to calculate the total weighted budget value. S66. Divide the weighted budget value of each coding tree unit by the total weighted budget value and multiply by the corrected bit budget to calculate the number of redistributed bits for each coding tree unit. Add the number of redistributed bits for all coding tree units to replace the original corrected bit budget, complete the second correction of the corrected bit budget and output it.

[0015] The beneficial effects of this invention are: This invention addresses the challenges of large differences in visually sensitive regions and high local texture complexity in video coding by constructing a decoupled structure for downsampling and spatiotemporal features. It generates spatially perceptual feature maps and temporally motion complexity feature maps through block variance calculation and motion amplitude extraction, respectively, and combines cross-modal attention fusion to output a dual-stream heterogeneous representation tensor. The dual-stream heterogeneous representation tensor is concatenated with pre-set state parameters and input into a nonlinear mapping network to mine the correlation between local texture and global bitrate budget, outputting dynamic adjustment coefficients for Lagrange multipliers. These coefficients are then used to weight and correct the baseline Lagrange multipliers to generate target Lagrange multipliers, which are then injected into the rate-distortion optimization layer of the video encoder for mode traversal and secondary solving, indirectly deriving the target quantization parameters for each coding tree unit. In practical coding... After coding, the instantaneous coding error is calculated, and the initial buffer fullness is input into the improved NMPC algorithm. By constructing a dual-branch heterogeneous prediction structure and introducing a hard-gated switching mechanism based on motion bursts, channel hard cut-off and activation-only operations are performed using the variance of the motion vector field as the criterion. The activated branch is used to infer the future error trajectory, construct and solve a rolling optimization problem with hard constraints to generate a candidate compensation sequence, and extract the first element as the closed-loop residual compensation bias output for the next frame. Furthermore, through a virtual buffer verification gating mechanism, the bias is superimposed on the bit budget reference surface to generate a corrected bit budget. If the initial buffer fullness is detected to exceed the safe threshold range, an adaptive bit forced reallocation algorithm based on visual perception weights is triggered to perform a second correction on the corrected bit budget. Ultimately, this achieves deep perception of the time-space characteristics in the video coding pipeline, precise intervention of the underlying coding parameters, agile response to sudden motion scenarios, and safe bottom-line closed-loop control of the buffer level, effectively improving the perceived rationality of bitrate resource allocation, the anti-divergence robustness of error trajectory prediction, and the system's anti-out-of-bounds stability under extreme budget offsets. Attached Figure Description

[0016] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings: Figure 1 This is a flowchart of an efficient image and video coding rate control method based on deep learning proposed in this invention; Figure 2 This is a flowchart of the spatiotemporal feature decoupling and dual-stream heterogeneous characterization tensor generation proposed in this invention; Figure 3 This is a flowchart of the error prediction and closed-loop compensation based on sudden motion hard-gated switching proposed in this invention. Detailed Implementation

[0017] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, illustrating only the basic structure of the invention, and therefore only show the components relevant to the invention.

[0018] refer to Figures 1-3 A high-efficiency image and video coding rate control method based on deep learning includes the following steps: S1. Acquire the current image frame to be encoded and the preceding state parameters in the video encoding pipeline. The preceding state parameters include the motion vector field distribution generated by forward motion estimation, the initial buffer fill degree, and the benchmark Lagrange multiplier calculated based on the traditional rate distortion model. Output the current image frame and the preceding state parameters. S2. Perform downsampling and spatiotemporal feature decoupling processing on the current image frame to construct a spatial perception feature map reflecting human visual sensitivity and a temporal motion complexity feature map reflecting coding difficulty, respectively. Perform cross-modal attention fusion and output a dual-stream heterogeneous representation tensor. S3. The dual-stream heterogeneous representation tensor is concatenated with the previous state parameters. The correlation between local texture and global bit rate budget in the current image frame is mined through nonlinear mapping. The Lagrange multiplier dynamic adjustment coefficients of each coding tree unit in the current image frame are output. S4. The reference Lagrange multipliers in the pre-state parameters are weighted and corrected using the Lagrange multipliers dynamic adjustment coefficients to generate the target Lagrange multipliers, which are then injected into the rate-distortion optimization layer of the video encoder for secondary solution, thereby indirectly deriving and outputting the target quantization parameters of each coding tree unit. S5. Calculate the instantaneous coding error after actual encoding based on the target quantization parameters, and input the improved NMPC algorithm with the initial buffer fullness input to construct a dual-branch heterogeneous prediction structure and introduce a hard-gating switching mechanism based on motion bursts. Use the variance of the motion vector field distribution as the criterion to perform hard cut-off and activation-only operations on the two prediction branches. Use the activated branches to infer the future error trajectory, construct and solve the rolling optimization problem with hard constraints to generate a candidate compensation sequence, and extract only the first element of the candidate compensation sequence as the closed-loop residual compensation bias output for the next frame. S6. The closed-loop residual compensation bias is superimposed on the bit budget reference plane of the next image frame to be encoded to generate the corrected bit budget. A virtual buffer verification gating mechanism is introduced. If the initial buffer fullness is detected to exceed the safe threshold range, an adaptive bit forced reallocation algorithm based on visual perception weight is triggered to perform a second correction on the corrected bit budget, thus completing efficient closed-loop rate control.

[0019] This invention significantly improves the perceptual rationality and system robustness of video coding rate control. By downsampling the current image frame and decoupling its spatiotemporal features, spatial and temporal feature maps are constructed and fused across modalities, achieving precise unified modeling of underlying visual sensitivity and coding difficulty, thus enhancing the adaptive expression capability for complex textures. Utilizing fused features to dynamically correct Lagrange multipliers and inject rate-distortion optimization layers, coding resources can be precisely tilted to areas sensitive to human vision, effectively avoiding localized mosaic effects and image quality collapse that occur in traditional models during intense motion. Introducing an improved NMPC algorithm, a hard-gating switching mechanism based on motion bursts is used to perform real-time hard severing and activation-only operations on the dual-branch prediction structure, achieving advanced prediction and compensation for future buffer overflows. Combining virtual buffer verification gating and an adaptive bit-forced reallocation algorithm for secondary correction completely eliminates the risk of buffer overflows caused by sudden bitrate changes. This method demonstrates strong anti-divergence capabilities when dealing with extreme situations such as rapid camera panning and sudden scene changes, and significantly improves the precision, smoothness, and intelligent level of closed-loop control in high dynamic video stream bitrate allocation.

[0020] In this embodiment, S1 specifically includes: S11. Extract the pixel displacement data generated in the forward motion estimation process in the video coding pipeline, arrange them into a two-dimensional matrix according to the original image resolution, and generate and output the motion vector field distribution. S12. Read the value of the current remaining bits in the buffer and the total capacity of the buffer during the video encoding pipeline operation, divide the current remaining bits by the total capacity of the buffer to calculate the ratio, generate and output the initial buffer fullness; S13. Divide the preset target bit rate value by the video frame rate value to calculate the target bit value of a single frame. Obtain the difference between the original pixel value and the reconstructed pixel value of all pixel blocks in the current image frame under the initial quantization parameters. Square the difference and add them to obtain the absolute error sum of squares value. Divide the target bit rate value of a single frame by the absolute error sum of squares value to obtain the slope value. Set the slope value as the reference Lagrange multiplier and output it. The preset target bit rate value is 2,000,000 bits per second and the video frame rate value is 30 frames per second. S14. Combine the pixel matrix, motion vector field distribution, initial buffer fill value, and reference Lagrange multiplier value of the current image frame to be encoded into the same data set in order, and output the current image frame and the previous state parameters.

[0021] In this embodiment, S2 specifically includes: S21. Input the pixel matrix of the current image frame into a downsampling convolutional layer of a preset size, add the gray values ​​of every two adjacent pixels in the horizontal direction and calculate the average value, add the gray values ​​of every two adjacent pixels in the vertical direction and calculate the average value, and generate a downsampling image reduced to the original preset ratio size, wherein the preset size is 3×3 and the preset ratio size is one-quarter. S22. Divide the downsampled image into multiple image blocks according to a preset pixel size. Subtract the average gray value of all pixels in each image block from the gray value of the image block and square the result. Add the squared results of all pixels and divide by the total number of pixels to calculate the variance value. Arrange the variance values ​​of all image blocks into a two-dimensional matrix in row and column order to generate a spatial perception feature map reflecting the visual sensitivity of the human eye. The preset pixel size is 16×16. S23. Divide the current image frame into multiple image blocks according to the preset pixel size, extract the horizontal displacement value and vertical displacement value of the image block corresponding to the motion vector field distribution in the previous state parameters, add the horizontal and vertical displacement values ​​of all pixels in the same image block and take the absolute value, and add them to obtain the motion amplitude value. Arrange the motion amplitude values ​​of all image blocks into a two-dimensional matrix in the row and column order to generate a temporal motion complexity feature map that reflects the coding difficulty. S24. Multiply the variance value of each image block in the spatial perception feature map with the motion amplitude value of the corresponding image block in the temporal motion complexity feature map to obtain the correlation value. Divide each variance value in the spatial perception feature map by the sum of all variance values ​​to obtain the first weight. Divide each motion amplitude value in the temporal motion complexity feature map by the sum of all motion amplitude values ​​to obtain the second weight. S25. Multiply the associated values ​​by the product of the first weight and the second weight, arrange all the multiplication results in their original positions and add them together to generate and output the dual-stream heterogeneous representation tensor.

[0022] In this embodiment, S3 specifically includes: S31. Expand the dual-stream heterogeneous representation tensor into a one-dimensional feature sequence along the channel dimension, extract the initial buffer fullness and the benchmark Lagrange multiplier from the pre-state parameters, and multiply them by the corresponding preset scaling factors respectively. Then, append the two scaled values ​​to the end of the one-dimensional feature sequence in sequence to generate a global state enhancement feature vector. The preset scaling factors are 100 and 0.01. S32. Input the global state enhancement feature vector into the first reshaping layer of the nonlinear mapping network, rearrange the global state enhancement feature vector into a three-dimensional tensor, divide the three-dimensional tensor into multiple local windows along the depth dimension, calculate the local bias by subtracting the average value of all values ​​in the window from the value at each position in each local window, add the local bias to the original value to output the local texture self-calibration tensor, the dimension of the three-dimensional tensor is 256×8×8, the size of the local window is 8×8×8 and the number is 16; S33. Input the local texture self-calibration tensor into the cross-attention layer of the nonlinear mapping network, divide the local texture self-calibration tensor into a query matrix, a key matrix, and a value matrix, transpose the query matrix and the key matrix, multiply them, and divide by a preset square root value to obtain a similarity weight matrix, input the value of each row in the similarity weight matrix into the Softmax function to calculate the exponential ratio, multiply the exponential ratio with the value matrix and sum over the depth dimension to output a local global association tensor. The dimensions of the query matrix, key matrix, and value matrix are 256×64, 64×256, and 64×256, respectively, and the preset square root value is 8. S34. Input the local-global correlation tensor into the two one-dimensional convolutional layers of the feedforward network layer of the nonlinear mapping network. Between the two one-dimensional convolutional layers, divide each value by a preset ratio value and add the corresponding value to calculate the normalized value. After the second one-dimensional convolutional layer, perform a Gaussian error linear unit activation operation on each value to map negative values ​​to output values ​​that decay exponentially and output the nonlinear mapping tensor. The number of output channels of the two one-dimensional convolutional layers are 512 and 256, respectively, and the preset ratio value is 512. S35. Input the nonlinear mapping tensor into the output layer of the nonlinear mapping network, and process it sequentially through one-dimensional convolution operation and Sigmoid activation function to generate a continuous floating-point value sequence. Divide the current image frame into multiple grid regions according to the preset division size of the coding tree unit. Assign the floating-point values ​​in the continuous floating-point value sequence to the corresponding grid regions as Lagrange multiplier dynamic adjustment coefficients and output them. The preset division size is 64×64 pixels.

[0023] In this embodiment, S4 specifically includes: S41. Extract the dynamic adjustment coefficients of the Lagrange multipliers corresponding to each coding tree unit in the current image frame, read the reference Lagrange multipliers corresponding to the same coding tree unit in the pre-state parameters, multiply the dynamic adjustment coefficients of the Lagrange multipliers directly with the reference Lagrange multipliers, generate and output the target Lagrange multipliers of the corresponding coding tree unit. S42. Input the target Lagrange multiplier of each coding tree unit into the rate-distortion optimization layer of the video encoder. In the rate-distortion optimization layer, obtain the sum of squared absolute errors between the original pixel value and the predicted pixel value of the corresponding coding tree unit in the current prediction mode. Multiply the sum of squared absolute errors by the target Lagrange multiplier to obtain the distortion cost value. S43. For the prediction mode index, motion vector data and residual matrix data generated by the corresponding coding tree unit in the current prediction mode, convert the prediction mode index, motion vector data and residual matrix data into binary bit streams in sequence, count the total number of bits contained in the binary bit stream as the actual bit value of encoding consumption, and directly add the distortion cost value to the actual bit value of encoding consumption to calculate the total distortion cost value. S44. Traverse all selectable prediction modes of the corresponding coding tree unit and repeat the step of calculating the total rate distortion cost value a preset number of times. Compare all total rate distortion cost values, find the optimal prediction mode corresponding to the total rate distortion cost value with the smallest value, and extract the residual data under the optimal prediction mode. The preset number of times is 9. S45. Input the residual data in the optimal prediction mode into the quantizer, divide the value of the residual data by the initial quantization step size and round down to obtain the integer coefficient, multiply it by the initial quantization step size to obtain the quantized residual value, and calculate the difference between the value of the residual data and the quantized residual value as the quantization error value. S46. Multiply the quantization error value by a preset compensation coefficient to obtain a compensation offset value, subtract the compensation offset value from the initial quantization step size to obtain a target quantization step size, divide the target quantization step size by a preset logarithmic base to obtain a logarithmic value, and add the logarithmic value to a preset offset constant to obtain a target quantization parameter. The preset compensation coefficient is 0.85, the preset logarithmic base is 4.7, and the preset offset constant is 6.

[0024] In this embodiment, S5 specifically includes: S51. Extract the single-frame target bit value allocated in the current image frame and the actual number of bits consumed after actual encoding. Subtract the actual number of bits consumed from the single-frame target bit value to calculate the difference and generate the instantaneous encoding error. S52. Extract the initial buffer fullness from the pre-state parameters, concatenate the instantaneous coding error value with the initial buffer fullness value in the order of front and back to generate a two-dimensional state vector, input the two-dimensional state vector into the improved NMPC algorithm, construct a dual-branch heterogeneous prediction structure inside the improved NMPC algorithm, set up the first calculation channel containing two one-dimensional convolutional layers as the stationary prediction branch, and set up the second calculation channel containing three fully connected layers with the number of output nodes decreasing sequentially as the mutation prediction branch. S53. Extract the motion vector field distribution from the previous state parameters. Subtract the corresponding average horizontal displacement and average vertical displacement from the horizontal displacement values ​​and vertical displacement values ​​of all pixels in the motion vector field distribution. Square the results of the subtraction and sum them up. Then divide by the total number of pixels in the motion vector field distribution to calculate the variance. S54. The calculated variance is compared with the preset mutation threshold as the motion burst criterion. A hard gating switching mechanism is executed. When the variance is greater than or equal to the preset mutation threshold, it is determined to be a motion burst. The values ​​of all positions in the two one-dimensional convolutional layers in the first calculation channel are directly replaced with 0 to perform a hard cut-off operation. The numerical calculation process of the three fully connected layers in the second calculation channel is retained. Only the mutation prediction branch is activated. The preset mutation threshold is 50. S55. When the variance is less than the preset mutation threshold, the motion is judged to be stable. The values ​​of all positions in the three fully connected layers in the second calculation channel are directly replaced with 0 to perform a hard cut-off operation, while the numerical calculation process of the two one-dimensional convolutional layers in the first calculation channel is retained, and only the stable prediction branch is activated. S56. Generate future error trajectories in the only activated stationary prediction branch or abrupt prediction branch, and extract the number of compensation bits ranked first in the calculated candidate compensation sequence. Use the extracted number of compensation bits directly as the closed-loop residual compensation bias of the next frame and output it.

[0025] In this embodiment, S56 specifically includes: S561. In the stationary prediction branch that is activated only, the two-dimensional state vector is input into the first one-dimensional convolutional layer. The value of the one-dimensional convolutional kernel at each position is multiplied and added with the value at the corresponding position in the two-dimensional state vector. After passing through the ReLU activation function, the one-dimensional intermediate vector is output and input into the second one-dimensional convolutional layer to repeat the operation of multiplication, addition and replacement of negative numbers, and output the predicted value at the current time. S562. The current predicted value plus the instantaneous coding error is used as the two-dimensional state vector of the next moment and input into two one-dimensional convolutional layers. The summation input operation is performed continuously for a preset number of times. The predicted values ​​output each time are arranged in chronological order to generate the future error trajectory. The preset number of times is 5. S563. In the mutation prediction branch that is only activated, the two-dimensional state vector is input into the first fully connected layer for fully connected layer linear mapping. The first bias value is added and the calculation result is compressed to between -1 and 1 by the hyperbolic tangent function to output the first hidden vector. After inputting the second fully connected layer to output the second hidden vector, the second hidden vector is input into the third fully connected layer to output the predicted value at the current time. S564. The current predicted value plus the instantaneous coding error is used as the two-dimensional state vector of the next moment and input into three fully connected layers. The summation input operation is performed continuously a preset number of times. The predicted values ​​output each time are arranged in chronological order to generate the future error trajectory. The preset number of times is 5. S565. In the improved NMPC algorithm, the upper limit and lower limit of the buffer fullness corresponding to each future image frame are preset as the constraint boundary of the rolling optimization problem with hard constraints. For a preset number of predicted values ​​in the future error trajectory, the same preset number of compensation bits to be calculated are initialized with an initial value of 0. The future buffer prediction value is calculated by subtracting the corresponding compensation bit number from each predicted value. The upper limit of the buffer fullness is 0.8, the lower limit of the buffer fullness is 0.2, and the preset number is 5. S566. Determine whether each future buffer prediction value is between the upper limit of buffer fullness and the lower limit of buffer fullness. When it exceeds the value range, calculate the difference between the future buffer prediction value and the upper limit of buffer fullness or the lower limit of buffer fullness as the out-of-bounds deviation value. Multiply the out-of-bounds deviation value by a preset penalty weight to obtain the penalty cost value. The preset penalty weight is 1000. S567. The total cost value is calculated by adding the sum of the squares of the preset number of compensation bits to all penalty cost values. The value of the preset number of compensation bits is changed one by one within the preset search step size and the total cost value is calculated repeatedly. All calculation results are compared, and the 5 compensation bits that minimize the total cost value are extracted and arranged in chronological order to generate a candidate compensation sequence. The preset search step size is 100 bits. S568. In the generated candidate compensation sequence, extract the number of compensation bits ranked first, and directly use the extracted number of compensation bits as the closed-loop residual compensation bias of the next frame and output it.

[0026] This invention introduces an improved NMPC algorithm combined with a dual-branch heterogeneous prediction structure to achieve advanced prediction and accurate compensation of video coding buffer levels. Instantaneous coding error and initial buffer fullness are constructed as a two-dimensional state vector, establishing a stationary prediction branch based on one-dimensional convolution and a sudden prediction branch based on a fully connected layer. A hard-gating switching mechanism is executed using the variance of the motion vector field as the criterion. When the variance reaches a preset threshold, the parameters of idle branches are directly set to zero and hard-cut off, activating only the corresponding branch to predict future error trajectories. Based on this, a hard-constrained rolling optimization problem is constructed, using the upper and lower limits of the buffer as boundaries. Candidate compensation sequences are solved by minimizing the total cost function including out-of-bounds penalties. This invention can completely solve the buffer overflow or exhaustion problem caused by sudden motion changes in traditional rate control, extracting the first-order compensation value to achieve single-step closed-loop intervention, significantly suppressing bitrate jumps, and exhibiting strong anti-divergence robustness in extreme scenarios such as drastic shot transitions.

[0027] The improved NMPC algorithm of this invention is similar to the traditional NMPC algorithm in that both retain the core architecture of model predictive control, that is, both construct the future state prediction trajectory in the finite time domain based on the current system state, both use a constrained rolling optimization mechanism to solve the optimal control sequence, and both follow the rolling time domain control principle of applying only the first control variable in the solution sequence to the current controlled object.

[0028] The difference lies in that this invention breaks away from the limitation of traditional NMPC algorithms relying on a single fixed mathematical model for state prediction and optimization, introducing a dual-branch heterogeneous prediction structure and a hard-gated switching mechanism. Building upon traditional NMPC's direct deduction of state trajectories using mechanistic models, this invention pre-lays a burst discrimination module based on the variance of the motion vector field. This module, acting as an independent routing mechanism, can assess the intensity of motion in the image in real time. When the variance is greater than or equal to a preset threshold, a hard cut-off is performed by directly setting the network parameters of idle branches to zero, forcing the model to activate only the abrupt prediction branch containing fully connected layers at the current moment; conversely, only the stationary prediction branch containing one-dimensional convolutional layers is activated. Subsequently, the activated branches independently deduce future error trajectories using their dedicated network structures. In the hard-constrained rolling optimization, a penalty cost numerical calculation mechanism based on multiplying the out-of-bounds deviation by a preset penalty weight is introduced to obtain candidate compensation sequences.

[0029] Based on the above improvements, the beneficial effects of this invention are that, through hard-gated switching and a dual-branch proprietary network, the improved NMPC algorithm can adaptively select the matching prediction mode for drastically different motion states in the video stream. This breaks the limitation of traditional NMPC, which uses a single fixed model and cannot balance prediction accuracy and computation speed, and realizes the dynamic and heterogeneous nature of the error inference strategy. The introduced high-weight penalty cost mechanism ensures absolute respect for the safety boundary of the buffer during the optimization process, effectively preventing the water level from exceeding the boundary. The hard cut-off mechanism improves the accuracy of advance prediction in sudden change scenarios while directly shielding the invalid calculation of redundant branches, significantly enhancing the anti-divergence robustness and real-time response capability of the video coding rate control link in extreme scenarios such as drastic shot switching.

[0030] In this embodiment, S6 specifically includes: S61. Extract the single-frame target bit value corresponding to the current image frame as the bit budget reference surface for the next image frame to be encoded, read the closed-loop residual compensation bias, add the value of the bit budget reference surface to the value of the closed-loop residual compensation bias, and calculate and generate the corrected bit budget. S62. Extract the total number of coding tree units contained in the next image frame to be encoded, divide the corrected bit budget by the total number of coding tree units to calculate the initial allocated number of bits for each coding tree unit, subtract the corrected bit budget from the current remaining number of bits and add the actual number of bits consumed to calculate the predicted remaining number of bits, and divide the predicted remaining number of bits by the total capacity of the buffer to calculate the prediction buffer fullness. S63. The upper and lower limits of the preset safety threshold range are compared with the fullness of the prediction buffer. When the fullness of the prediction buffer is between the upper and lower limits, the corrected bit budget is directly output as the final bit budget to complete efficient closed-loop rate control. When the fullness of the prediction buffer is greater than the upper limit or less than the lower limit, the virtual buffer verification gating mechanism is triggered. The upper limit of the safety threshold range is 0.85 and the lower limit of the safety threshold range is 0.15. S64. When the virtual buffer verification gating mechanism is triggered, extract the spatial perception feature map corresponding to the current image frame, divide the variance value of each image block in the spatial perception feature map by the sum of the variance values ​​of all image blocks, and calculate and generate the visual perception weight corresponding to each coding tree unit. S65. Trigger the adaptive bit forced reallocation algorithm based on visual perception weights, multiply the visual perception weight of each coding tree unit by the corrected bit budget to calculate the weighted budget value, and add the weighted budget values ​​of all coding tree units to calculate the total weighted budget value. S66. Divide the weighted budget value of each coding tree unit by the total weighted budget value and multiply by the corrected bit budget to calculate the number of redistributed bits for each coding tree unit. Add the number of redistributed bits for all coding tree units to replace the original corrected bit budget, complete the second correction of the corrected bit budget and output it.

[0031] Example 1: To verify the feasibility of this invention in practice, it was applied to an ultra-high-definition video transcoding and distribution platform of a large-scale integrated media cultural communication center in a certain city. This platform primarily undertakes the transcoding of 4K and 8K ultra-high-definition video live streams for various large-scale sports events, outdoor music festivals, and important government activities throughout the city. It processes over forty live video streams concurrently daily, with the average bitrate of a single stream fluctuating between 15 and 20 Mbps. In actual live streaming scenarios, rapid camera panning, instantaneous close-up shots, and complex nighttime lighting changes occur extremely frequently. This highly dynamic video content poses a significant challenge to the bit allocation of the underlying encoder. Traditional transcoding platforms generally rely on static lookup tables and fixed-rate distortion models, which cannot adjust quantization parameters in a timely manner when faced with sudden motion scenarios. This often results in instantaneous blurring, a surge in pixelation, and severe fluctuations in end-to-end live streaming latency, seriously affecting the viewing experience of millions of viewers. It also easily leads to buffer overflows in downstream distribution nodes, causing stuttering and interrupted streaming.

[0032] In practical deployment, the method of this invention is seamlessly embedded into the video encoding control pipeline of the platform, replacing the original fixed bitrate control module. Before entering the encoding kernel, the front-end video sources connected to the platform each day undergo downsampling and block processing by the feature perception module of this invention. This extracts spatial perception features reflecting texture details and temporal motion complexity features reflecting the dynamic state of the image, which are then fused to generate a dual-stream heterogeneous representation tensor. This tensor is directly used to dynamically correct the Lagrange multipliers, changing the mode decision process of the encoding tree unit from the source, allowing smooth areas sensitive to the human eye and areas of violent motion to receive drastically different bit resource tilts. After encoding is completed and the actual consumed bits are generated, the system synchronously extracts the initial buffer fullness and instantaneous encoding error, directly inputting them into the improved NMPC algorithm. At this point, the dual-branch heterogeneous prediction structure built within the algorithm begins to function. The system calculates the variance of the motion vector field of the current frame in real time. Once the variance is detected to exceed the preset mutation threshold, the system will decisively trigger a hard-gating switching mechanism, instantly cutting off the steady prediction branch responsible for normal, smooth frames and activating only the mutation prediction branch containing a deep fully connected network. The activated branch will combine historical error trajectories to deduce the potential risk of buffer level overflow in the future, and search for the optimal number of compensation bits in a rolling optimization solver with hard constraints. Finally, the first compensation element extracted is added as a bias to the baseline bit budget of the next frame, thus completing the buffer level overflow prevention intervention at the moment of motion mutation. Through this link from feature perception to advanced prediction to closed-loop compensation, the platform achieves flexible and precise bitrate control for extremely unstable video streams. During the two-month trial operation, the method of this invention and the platform's original bitrate control method based on fixed lookup tables and simple quadratic models were compared in several typical high-dynamic live streaming scenarios. Table 1 below shows the actual operation comparison data extracted during this period: Table 1. Comparison of bitrate control and image quality protection performance between the present invention and traditional methods.

[0033] As can be clearly seen from the detailed comparison data shown in Table 1, the rate control method based on dual-branch heterogeneous prediction and improved NMPC algorithm proposed in this invention exhibits extremely significant performance advantages and strong environmental adaptability when dealing with various complex and ever-changing live streaming scenarios.

[0034] Regarding the stability of bitrate control, traditional methods, lacking the ability to predict future errors, generally exhibit average bitrate deviations hovering at a high level of 10% to 18%. This is especially true in nighttime live streaming scenarios with frequent light flickering, where the deviation can even reach 18.2%. This means the actual bitrate delivered deviates significantly from the target, easily overwhelming downstream network nodes. In contrast, this invention utilizes an improved NMPC algorithm for proactive rolling optimization, firmly suppressing the average bitrate deviation across all scenarios to within 3.5%. In demanding scenarios like government event filming, the deviation is further reduced to 1.5%, ensuring precise utilization of network bandwidth resources.

[0035] In protecting image quality in the face of sudden extreme motion events, the core advantages of this invention are fully demonstrated. Traditional methods, in fast-moving scenes such as sporting events and music festivals, often result in a maximum jump variable exceeding two megabits per frame, causing frequent buffer overflows and forcibly requantizing the encoder. This directly manifests as severe mosaic bursts on the screen; in one instance, nighttime scenes experienced forty-two severe image quality collapses within two weeks. In contrast, this invention introduces a hard-gating switching mechanism based on sudden motion events. Upon detecting a dramatic change in the motion vector, it immediately switches to the mutation prediction branch and outputs a closed-loop compensation bias, smoothly suppressing the maximum jump variable per frame to the hundreds of kilobits. In hundreds of hours of various tests, this invention, except for one minor image quality fluctuation triggered under extreme lighting, achieved a breakthrough of zero severe mosaic bursts in scenarios that put the system to its limits, such as sporting events and music festivals, completely eliminating sudden image degradation.

[0036] In terms of buffer level control and live streaming latency optimization, traditional methods, due to their passive handling of errors, result in a persistently high number of buffer overflows. This forces the system to add additional buffer queues as a buffer buffer, leading to an end-to-end average latency consistently exceeding 1,000 milliseconds, and reaching as high as 1,520 milliseconds in nighttime scenarios, completely failing to meet the low-latency requirements of interactive live streaming. This invention, by achieving accurate early prediction of buffer levels and hard constraint fallback, reduces the number of overflows from hundreds to single digits, eliminating redundant queuing time introduced by overflow prevention. This results in an end-to-end average latency consistently within an extremely low range of 800 to 920 milliseconds across all scenarios, achieving a speedup of nearly 40% compared to traditional methods. Overall, this invention not only completely solves the problems of bitrate runaway and image quality collapse in ultra-high-definition video streams under complex dynamic environments, but also provides a highly reliable and efficient underlying control paradigm for next-generation low-latency, high-quality interactive live streaming distribution networks.

[0037] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.

Claims

1. A high-efficiency image and video coding rate control method based on deep learning, characterized in that, Includes the following steps: S1. Acquire the current image frame to be encoded and the previous state parameters in the video encoding pipeline, and output the current image frame and the previous state parameters. S2. Perform downsampling and spatiotemporal feature decoupling processing on the current image frame to construct a spatial perception feature map reflecting human visual sensitivity and a temporal motion complexity feature map reflecting coding difficulty, respectively. Perform cross-modal attention fusion and output a dual-stream heterogeneous representation tensor. S3. The dual-stream heterogeneous representation tensor is concatenated with the previous state parameters. The correlation between local texture and global bit rate budget in the current image frame is mined through nonlinear mapping, and the dynamic adjustment coefficient of the Lagrange multiplier is output. S4. The reference Lagrange multipliers in the pre-state parameters are weighted and corrected using the Lagrange multipliers dynamic adjustment coefficients to generate the target Lagrange multipliers, which are then injected into the rate-distortion optimization layer of the video encoder for secondary solution, thereby indirectly deriving and outputting the target quantization parameters of each coding tree unit. S5. Calculate the instantaneous coding error after actual coding based on the target quantization parameters, and input the improved NMPC algorithm with the initial buffer fullness to construct a dual-branch heterogeneous prediction structure and introduce a hard-gating switching mechanism based on motion bursts. Use the variance of the motion vector field distribution as the criterion to perform hard cut-off and activation-only operations on the two prediction branches, deduce the future error trajectory, generate candidate compensation sequences, and output the closed-loop residual compensation bias. S6. The closed-loop residual compensation bias is superimposed on the bit budget reference plane of the next image frame to be encoded to generate the corrected bit budget. If the initial buffer fullness is detected to exceed the safe threshold range, the adaptive bit forced reallocation algorithm is triggered to perform secondary correction and complete efficient closed-loop rate control.

2. The efficient image and video coding rate control method based on deep learning according to claim 1, characterized in that, S1 specifically includes: S11. Extract the pixel displacement data generated in the forward motion estimation process in the video coding pipeline, arrange them into a two-dimensional matrix according to the original image resolution, and generate and output the motion vector field distribution. S12. Read the value of the current remaining bits in the buffer and the total capacity of the buffer during the video encoding pipeline operation, divide the current remaining bits by the total capacity of the buffer to calculate the ratio, generate and output the initial buffer fullness; S13. Divide the preset target bit rate value by the video frame rate value to calculate the target bit value of a single frame. Obtain the difference between the original pixel value and the reconstructed pixel value of all pixel blocks in the current image frame under the initial quantization parameters. Square the difference and add them together to obtain the absolute error sum of squares value. Divide the target bit value of a single frame by the absolute error sum of squares value to obtain the slope value. Set the slope value as the reference Lagrange multiplier and output it. S14. Combine the pixel matrix, motion vector field distribution, initial buffer fill value, and reference Lagrange multiplier value of the current image frame to be encoded into the same data set in order, and output the current image frame and the previous state parameters.

3. The efficient image and video coding rate control method based on deep learning according to claim 1, characterized in that, S2 specifically includes: S21. Input the pixel matrix of the current image frame into a downsampling convolutional layer of a preset size, add the gray values ​​of every two adjacent pixels in the horizontal direction and calculate the average value, add the gray values ​​of every two adjacent pixels in the vertical direction and calculate the average value, and generate a downsampling image reduced to the original preset ratio size. S22. Divide the downsampled image into multiple image blocks according to a preset pixel size. Subtract the average gray value of all pixels in each image block from the average gray value of the image block and square the result. Add the squared results of all pixels and divide by the total number of pixels to calculate the variance value. Arrange the variance values ​​of all image blocks into a two-dimensional matrix in row and column order to generate a spatial perception feature map. S23. Divide the current image frame into multiple image blocks according to the preset pixel size, extract the horizontal displacement value and vertical displacement value of the image block corresponding to the motion vector field distribution in the previous state parameters, add the horizontal and vertical displacement values ​​of all pixels in the same image block and take the absolute value, and add them to obtain the motion amplitude value. Arrange the motion amplitude values ​​of all image blocks into a two-dimensional matrix in the row and column order to generate a temporal motion complexity feature map. S24. Multiply the variance value of each image block in the spatial perception feature map with the motion amplitude value of the corresponding image block in the temporal motion complexity feature map to obtain the correlation value. Divide each variance value in the spatial perception feature map by the sum of all variance values ​​to obtain the first weight. Divide each motion amplitude value in the temporal motion complexity feature map by the sum of all motion amplitude values ​​to obtain the second weight. S25. Multiply the associated values ​​by the product of the first weight and the second weight, arrange all the multiplication results in their original positions and add them together to generate and output the dual-stream heterogeneous representation tensor.

4. The efficient image and video coding rate control method based on deep learning according to claim 1, characterized in that, S3 specifically includes: S31. Expand the dual-stream heterogeneous representation tensor into a one-dimensional feature sequence along the channel dimension, extract the initial buffer fullness and the benchmark Lagrange multiplier from the pre-state parameters, and multiply them by the corresponding preset scaling factor respectively. Then, append the two scaled values ​​to the end of the one-dimensional feature sequence in sequence to generate a global state enhancement feature vector. S32. Input the global state enhancement feature vector into the first reshaping layer of the nonlinear mapping network, rearrange the global state enhancement feature vector into a three-dimensional tensor, divide the three-dimensional tensor into multiple local windows along the depth dimension, calculate the local bias by subtracting the average value of all values ​​in the window from the value at each position in each local window, and add the local bias to the original value to output the local texture self-calibration tensor. S33. Input the local texture self-calibration tensor into the cross-attention layer of the nonlinear mapping network, divide the local texture self-calibration tensor into a query matrix, a key matrix and a value matrix, transpose the query matrix and the key matrix, multiply them and divide by the preset square root value to obtain the similarity weight matrix, input the value of each row in the similarity weight matrix into the Softmax function to calculate the exponential ratio, multiply the exponential ratio with the value matrix and sum over the depth dimension, and output the local global association tensor. S34. Input the local-global correlation tensor into the two one-dimensional convolutional layers of the feedforward network layer of the nonlinear mapping network. Between the two one-dimensional convolutional layers, divide each value by a preset ratio and add the corresponding value to calculate the normalized value. After the second one-dimensional convolutional layer, perform Gaussian error linear unit activation operation on each value to map the negative value to an exponentially decaying output value and output the nonlinear mapping tensor. S35. Input the nonlinear mapping tensor into the output layer of the nonlinear mapping network, and process it sequentially through one-dimensional convolution operation and Sigmoid activation function to generate a continuous floating-point value sequence. Divide the current image frame into multiple grid regions according to the preset division size of the coding tree unit, and assign the floating-point values ​​in the continuous floating-point value sequence to the corresponding grid regions as Lagrange multiplier dynamic adjustment coefficients and output them.

5. The efficient image and video coding rate control method based on deep learning according to claim 1, characterized in that, S4 specifically includes: S41. Extract the dynamic adjustment coefficients of the Lagrange multipliers corresponding to each coding tree unit in the current image frame, read the reference Lagrange multipliers corresponding to the same coding tree unit in the pre-state parameters, multiply the dynamic adjustment coefficients of the Lagrange multipliers directly with the reference Lagrange multipliers, generate and output the target Lagrange multipliers of the corresponding coding tree unit. S42. Input the target Lagrange multiplier of each coding tree unit into the rate-distortion optimization layer of the video encoder. In the rate-distortion optimization layer, obtain the sum of squared absolute errors between the original pixel value and the predicted pixel value of the corresponding coding tree unit in the current prediction mode. Multiply the sum of squared absolute errors by the target Lagrange multiplier to obtain the distortion cost value. S43. For the prediction mode index, motion vector data and residual matrix data generated by the corresponding coding tree unit in the current prediction mode, convert the prediction mode index, motion vector data and residual matrix data into binary bit streams in sequence, count the total number of bits contained in the binary bit stream as the actual bit value of encoding consumption, and directly add the distortion cost value to the actual bit value of encoding consumption to calculate the total distortion cost value. S44. Traverse all possible prediction modes of the corresponding coding tree unit and repeat the step of calculating the total rate distortion cost value a preset number of times. Compare all total rate distortion cost values, find the optimal prediction mode corresponding to the total rate distortion cost value with the smallest value, and extract the residual data under the optimal prediction mode. S45. Input the residual data in the optimal prediction mode into the quantizer, divide the value of the residual data by the initial quantization step size and round down to obtain the integer coefficient, multiply it by the initial quantization step size to obtain the quantized residual value, and calculate the difference between the value of the residual data and the quantized residual value as the quantization error value. S46. Multiply the quantization error value by the preset compensation coefficient to obtain the compensation offset value, subtract the compensation offset value from the initial quantization step size to obtain the target quantization step size, divide the target quantization step size by the preset logarithmic base to obtain the logarithmic value, and add the logarithmic value to the preset offset constant to obtain the target quantization parameter.

6. The efficient image and video coding rate control method based on deep learning according to claim 1, characterized in that, S5 specifically includes: S51. Extract the single-frame target bit value allocated in the current image frame and the actual number of bits consumed after actual encoding. Subtract the actual number of bits consumed from the single-frame target bit value to calculate the difference and generate the instantaneous encoding error. S52. Extract the initial buffer fullness from the pre-state parameters, concatenate the instantaneous coding error value with the initial buffer fullness value in the order of front and back to generate a two-dimensional state vector, input the two-dimensional state vector into the improved NMPC algorithm, construct a dual-branch heterogeneous prediction structure inside the improved NMPC algorithm, set up the first calculation channel containing two one-dimensional convolutional layers as the stationary prediction branch, and set up the second calculation channel containing three fully connected layers with the number of output nodes decreasing sequentially as the mutation prediction branch. S53. Extract the motion vector field distribution from the previous state parameters. Subtract the corresponding average horizontal displacement and average vertical displacement from the horizontal displacement values ​​and vertical displacement values ​​of all pixels in the motion vector field distribution. Square the results of the subtraction and sum them up. Then divide by the total number of pixels in the motion vector field distribution to calculate the variance. S54. The calculated variance is compared with the preset mutation threshold as the motion burst criterion. A hard gating switching mechanism is executed. When the variance is greater than or equal to the preset mutation threshold, it is determined to be a motion burst. The values ​​of all positions in the two one-dimensional convolutional layers in the first calculation channel are directly replaced with 0 to perform a hard cut-off operation. The numerical calculation process of the three fully connected layers in the second calculation channel is retained, and only the mutation prediction branch is activated. S55. When the variance is less than the preset mutation threshold, the motion is judged to be stable. The values ​​of all positions in the three fully connected layers in the second calculation channel are directly replaced with 0 to perform a hard cut-off operation, while the numerical calculation process of the two one-dimensional convolutional layers in the first calculation channel is retained, and only the stable prediction branch is activated. S56. Generate future error trajectories in the only activated stationary prediction branch or abrupt prediction branch, and extract the number of compensation bits ranked first in the calculated candidate compensation sequence. Use the extracted number of compensation bits directly as the closed-loop residual compensation bias of the next frame and output it.

7. The efficient image and video coding rate control method based on deep learning according to claim 6, characterized in that, S56 specifically includes: S561. In the stationary prediction branch that is activated only, the two-dimensional state vector is input into the first one-dimensional convolutional layer. The value of the one-dimensional convolutional kernel at each position is multiplied and added with the value at the corresponding position in the two-dimensional state vector. After passing through the ReLU activation function, the one-dimensional intermediate vector is output and input into the second one-dimensional convolutional layer to repeat the operation of multiplication, addition and replacement of negative numbers, and output the predicted value at the current time. S562. The current predicted value plus the instantaneous coding error is used as the two-dimensional state vector of the next moment and input into two one-dimensional convolutional layers. The summation input operation is performed continuously a preset number of times. The predicted values ​​output each time are arranged in chronological order to generate the future error trajectory. S563. In the mutation prediction branch that is only activated, the two-dimensional state vector is input into the first fully connected layer for fully connected layer linear mapping, the first bias value is added, the hyperbolic tangent function is applied, and the calculation result is compressed to between -1 and 1 to output the first hidden vector. After inputting into the second fully connected layer to output the second hidden vector, the second hidden vector is input into the third fully connected layer to output the predicted value at the current time. S564. The current predicted value plus the instantaneous coding error is used as the two-dimensional state vector of the next moment and input into three fully connected layers. The summation input operation is performed continuously a preset number of times. The predicted values ​​output each time are arranged in chronological order to generate the future error trajectory. S565. In the improved NMPC algorithm, the upper limit and lower limit of the buffer fullness corresponding to each future image frame are preset as the constraint boundary of the rolling optimization problem with hard constraints. For the preset number of predicted values ​​in the future error trajectory, the same preset number of compensation bits to be obtained are initialized with an initial value of 0. The future buffer prediction value is calculated by subtracting the corresponding compensation bit number from each predicted value. S566. Determine whether each future buffer prediction value is between the upper limit of buffer fullness and the lower limit of buffer fullness. When it exceeds the value range, calculate the difference between the future buffer prediction value and the upper limit of buffer fullness or the lower limit of buffer fullness as the out-of-bounds deviation value. Multiply the out-of-bounds deviation value by the preset penalty weight to obtain the penalty cost value. S567. The total cost value is calculated by adding the sum of the squares of the preset number of compensation bits to all penalty cost values. Within the preset search step size, the value of the preset number of compensation bits is changed one by one and the total cost value is calculated repeatedly. All calculation results are compared, and the 5 compensation bits that minimize the total cost value are extracted and arranged in chronological order to generate a candidate compensation sequence. S568. In the generated candidate compensation sequence, extract the number of compensation bits ranked first, and directly use the extracted number of compensation bits as the closed-loop residual compensation bias of the next frame and output it.

8. The efficient image and video coding rate control method based on deep learning according to claim 1, characterized in that, S6 specifically includes: S61. Extract the single-frame target bit value corresponding to the current image frame as the bit budget reference surface for the next image frame to be encoded, read the closed-loop residual compensation bias, add the value of the bit budget reference surface to the value of the closed-loop residual compensation bias, and calculate and generate the corrected bit budget. S62. Extract the total number of coding tree units contained in the next image frame to be encoded, divide the corrected bit budget by the total number of coding tree units to calculate the initial allocated number of bits for each coding tree unit, subtract the corrected bit budget from the current remaining number of bits and add the actual number of bits consumed to calculate the predicted remaining number of bits, and divide the predicted remaining number of bits by the total capacity of the buffer to calculate the prediction buffer fullness. S63. The upper and lower limits of the preset safety threshold range are compared with the fullness of the prediction buffer and the safety threshold range. When the fullness of the prediction buffer is between the upper and lower limits, the corrected bit budget is directly used as the final bit budget output to complete efficient closed-loop rate control. When the fullness of the prediction buffer is greater than the upper limit or less than the lower limit, the virtual buffer verification gating mechanism is triggered. S64. When the virtual buffer verification gating mechanism is triggered, extract the spatial perception feature map corresponding to the current image frame, divide the variance value of each image block in the spatial perception feature map by the sum of the variance values ​​of all image blocks, and calculate and generate the visual perception weight corresponding to each coding tree unit. S65. Trigger the adaptive bit forced reallocation algorithm based on visual perception weights, multiply the visual perception weight of each coding tree unit by the corrected bit budget to calculate the weighted budget value, and add the weighted budget values ​​of all coding tree units to calculate the total weighted budget value. S66. Divide the weighted budget value of each coding tree unit by the total weighted budget value and multiply by the corrected bit budget to calculate the number of redistributed bits for each coding tree unit. Add the number of redistributed bits for all coding tree units to replace the original corrected bit budget, complete the second correction of the corrected bit budget and output it.