Video code rate control method and device, computer readable storage medium

By generating coding unit division information and quantization parameters through graph neural networks, the adaptability problem of video coding schemes when bandwidth changes is solved, coding efficiency and user experience are improved, and it is adaptable to multiple coding standards.

CN116320529BActive Publication Date: 2025-10-10SANECHIPS TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111508059.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-10
Publication Date
2025-10-10
Estimated Expiration
2041-12-10

AI Technical Summary

Technical Problem

Existing video coding solutions cannot provide adaptive bitrate transmission solutions when facing bandwidth changes, resulting in problems such as video freezes and ROI blur, and low coding efficiency.

Method used

Through graph neural networks, the global coding reference data of the compressed video is output with target constraints, the division information of the coding unit and the quantization parameters of the coding block are generated, the video coding efficiency is optimized, and it is adapted to multiple coding standards.

Benefits of technology

It improves video encoding efficiency, optimizes user viewing experience, reduces bit rate fluctuations, adapts to differences between different encoding standards and encoders, and avoids excessive degradation of video quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116320529B_ABST
    Figure CN116320529B_ABST
Patent Text Reader

Abstract

The application provides a video code rate control method and device, and a computer readable storage medium, wherein the method comprises: inputting global coding reference data of an acquired to-be-compressed video into a graph neural network to output code rate correlation data; and determining a current code rate parameter for controlling the video coding code rate according to the code rate correlation data; wherein the global coding reference data is used to represent the compression quality of the to-be-compressed video, and the code rate correlation data comprises at least one of the following types: division information of a coding unit, and a quantization parameter of each coding block in the coding unit. In the embodiment of the application, the current code rate parameter of the macro block level suitable for the application scene of the to-be-compressed video is obtained based on the code rate correlation data, which is beneficial to improving the video coding efficiency, optimizing the user viewing experience, and introducing no standard correlation information, and can better adapt to various coding standards.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] Embodiments of the present application relate to the technical field of video image processing, and particularly relate to a video code rate control method and device, and a computer readable storage medium. BACKGROUND

[0002] With the continuous development of network technology, device access requests and environments become complex and diverse, and gradually becoming an important issue is to overcome the experience decline caused by unstable bandwidth. Generally speaking, the most obvious influence of unstable bandwidth belongs to continuous traffic transmission, such as video signals; at present, the video coding scheme in the related art is relatively fixed, and is usually only applied to a specific coding standard, and when coping with a bandwidth change scene, it cannot provide an adaptive code rate transmission scheme for the same coding content, so the coding efficiency is relatively low, resulting in that users often have problems such as video stuttering, ROI (Region of Interest) picture blurring, or obvious subjective experience decline when watching videos. SUMMARY

[0003] The following is a summary of the subject matter of the detailed description herein. This summary is not intended to limit the scope of the claims.

[0004] Embodiments of the present application provide a video code rate control method and device, and a computer readable storage medium, which can improve video coding efficiency and optimize user viewing experience.

[0005] In a first aspect, embodiments of the present application provide a video code rate control method, comprising:

[0006] inputting the obtained global coding reference data of the to-be-compressed video into a graph neural network to output code rate associated data;

[0007] determining a current code rate parameter for controlling a video coding code rate according to the code rate associated data;

[0008] The global coding reference data is used to represent the compression quality of the to-be-compressed video, and the code rate associated data includes at least one of the following types: division information of a coding unit, and quantization parameters of each coding block in the coding unit.

[0009] In a second aspect, embodiments of the present application further provide a video code rate control device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, and the processor implements the video code rate control method of the first aspect when executing the computer program.

[0010] In a third aspect, embodiments of the present application further provide a computer readable storage medium, which stores computer executable instructions, and the computer executable instructions are used to execute the video code rate control method of the first aspect.

[0011] The embodiment of the present application comprises: inputting the obtained global coding reference data of the to-be-compressed video into a graph neural network to output code rate correlation data; determining a current code rate parameter for controlling the code rate of video coding according to the code rate correlation data; wherein the global coding reference data is used to represent the compression quality of the to-be-compressed video, and the code rate correlation data comprises at least one of the following types: division information of a coding unit, and quantization parameters of each coding block in the coding unit. According to the scheme provided by the embodiment of the present application, the global coding reference data of the to-be-compressed video is output by target constraint through the graph neural network, the code rate correlation data in the global optimization condition is obtained, the influence of code rate fluctuation caused by global error can be reduced, and then based on the division information of the coding unit or / and the quantization parameters of each coding block in the code rate correlation data, the current code rate parameter of the macroblock level suitable for the application scene of the to-be-compressed video is obtained, which is beneficial to improving the video coding efficiency, optimizing the user viewing experience, and does not introduce standard correlation information, and can better adapt to various coding standards.

[0012] Additional features and advantages of the application will be set forth in the description that follows, and in part will be apparent from the description, or can be learned by practice of the application. The objectives and other advantages of the application will be realized and attained by the structure particularly pointed out in the description and claims. BRIEF DESCRIPTION OF DRAWINGS

[0013] The accompanying drawings are included to provide a further understanding of the technical scheme of the application, and constitute a part of the specification, and are used together with the embodiments of the application to explain the technical scheme of the application, and do not constitute a limitation on the technical scheme of the application.

[0014] Figure 1 is a flowchart of a video code rate control method provided by an embodiment of the present application;

[0015] Figure 2 is a flowchart of determining a current code rate parameter in a video code rate control method provided by an embodiment of the present application;

[0016] Figure 3 is a structural schematic diagram of a graph neural network provided by an embodiment of the present application;

[0017] Figure 4 is a flowchart of outputting code rate correlation data in a video code rate control method provided by an embodiment of the present application;

[0018] Figure 5 is a flowchart of outputting code rate correlation data in a video code rate control method provided by another embodiment of the present application;

[0019] Figure 6is a flow chart of determining the first encoding quality evaluation parameter in the video code rate control method provided by one embodiment of the present application;

[0020] Figure 7 is a flow chart of determining the first encoding quality evaluation parameter in the video code rate control method provided by one embodiment of the present application;

[0021] Figure 8 is a flow chart of determining the encoding quality evaluation index corresponding to the reconstructed frame in the video code rate control method provided by one embodiment of the present application;

[0022] Figure 9 is an execution flow chart of determining the first encoding quality evaluation parameter;

[0023] Figure 10 is a flow chart of determining the encoding quality evaluation parameter in the video code rate control method provided by another embodiment of the present application;

[0024] Figure 11 is an execution flow chart of determining the second encoding quality evaluation parameter;

[0025] Figure 12 is a flow chart of obtaining the second code rate related data in the video code rate control method provided by one embodiment of the present application;

[0026] Figure 13 is a flow chart of obtaining the first code rate related data in the video code rate control method provided by one embodiment of the present application;

[0027] Figure 14 is a flow chart before outputting the code rate related data in the video code rate control method provided by one embodiment of the present application;

[0028] Figure 15 is a flow chart after determining the current code rate parameter in the video code rate control method provided by one embodiment of the present application;

[0029] Figure 16 is a schematic diagram of the video code rate control device provided by one embodiment of the present application. DETAILED DESCRIPTION

[0030] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application is further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and should not be used to limit the present application.

[0031] It should be noted that although the device schematics illustrate functional module divisions and the flowcharts illustrate logical sequences, in certain circumstances, the steps shown or described may be performed in a sequence that differs from the module divisions in the device or the sequence in the flowcharts. The terms "first," "second," and so on, used in the specification, claims, and drawings, are used to distinguish similar items and are not necessarily used to describe a specific sequence or precedence.

[0032] The present invention provides a video bit rate control method and device, and a computer-readable storage medium. Through a graph neural network, target constraints are output on the global coding reference data of the compressed video to obtain bit rate association data under global optimization, which can reduce the impact of bit rate fluctuations caused by global errors. Based on the division information of the coding units and / or the quantization parameters of each coding block in the bit rate association data, the current bit rate parameters at the macroblock level suitable for the application scenario of the video to be compressed are obtained, which is conducive to improving video coding efficiency and optimizing user viewing experience. It does not display the introduction of standard association information and can better adapt to multiple coding standards.

[0033] The embodiments of the present invention are further described below with reference to the accompanying drawings.

[0034] like Figure 1 As shown, Figure 1 4 is a flowchart of a video bit rate control method provided by an embodiment of the present invention. The video bit rate control method includes but is not limited to steps S100 to S200.

[0035] Step S100: Input the acquired global coding reference data of the video to be compressed into the graph neural network, and output bit rate associated data, wherein the global coding reference data is used to characterize the compression quality of the video to be compressed, and the bit rate associated data includes at least one of the following types: division information of the coding unit, and quantization parameters of each coding block in the coding unit.

[0036] In one embodiment, a graph neural network is used to perform target constraint output on the global coding reference data of the compressed video to obtain rate-related data under global optimization, which can reduce the impact of rate fluctuations caused by global errors, and the obtained rate-related data is the division information of the coding unit or / and the quantization parameters of each coding block in the coding unit. Those skilled in the art will know that the above two parameters are important indicators affecting video compression, so that relevant rate control parameters can be determined based on the division information of the coding unit or / and the quantization parameters of each coding block in the coding unit.

[0037] In one embodiment, the type of video to be compressed is not limited, and the method of obtaining the global coding reference data of the video to be compressed is not restricted and is well known to those skilled in the art and will not be described in detail here; the type of graph neural network (GNN) is not limited, and it can be a trained one. At this time, the global coding reference data is input into the trained graph neural network, and the trained graph neural network outputs the bit rate correlation data. The training method of the graph neural network is gradually explained in the following embodiments.

[0038] In one embodiment, global bitrate reference data is used to characterize the compression quality of a video to be compressed. Therefore, all factors that affect the compression quality of the video to be compressed may be considered as global bitrate reference data. In particular, unstructured data is highly independent and will not be affected by changes or modifications to other data, thus providing good reference value. For example, global bitrate reference data may include, but is not limited to, at least one of the following types:

[0039] Rate constraint information associated with the coding standard;

[0040] Region of Interest (ROI) information;

[0041] Encoding type information;

[0042] Encoder information;

[0043] Encoding frame constraint information;

[0044] Coded frame statistics;

[0045] Inter-frame information.

[0046] It should be noted that the rate constraint information associated with the coding standard may be pre-set, and there is corresponding rate constraint information for different coding standards.

[0047] It should be noted that the ROI information can be pre-set and used to characterize the encoding format supported by the encoder to determine whether the encoder supports ROI encoding. If such encoding strategy is supported, the ROI is prioritized, for example, initial values ​​from 0.1 to 1 are set according to the priority characteristics, 1 represents the highest priority, and 0.1 represents the lowest priority. Considering the convolution characteristics here, even the lowest priority is not described by a value of 0. On the contrary, if ROI encoding is not supported, the ROI matrix is ​​initialized to all 1s, thereby achieving a control strategy that supports whether or not a clear ROI is supported, greatly alleviating the excessive degradation of video quality in non-regions of interest (NROI), and reducing the volatility of the overall bit rate due to control deviation.

[0048] It should be noted that the encoding type information can cause different video compression scenarios, that is, it can affect the video code rate; the encoder information reflects the encoding influence of the encoder itself on the code rate related data, which can be caused by the structure, specification, etc. of the encoder itself, and needs to be analyzed and determined for specific encoders, which is not limited in the embodiment.

[0049] It should be noted that the encoding frame constraint information reflects the influence of the encoding frame information in the encoding process, which can be further determined based on the reference frame information, the current frame information, etc.

[0050] It should be noted that the encoding frame statistical information can be but not limited to texture information at the macroblock level, texture information of the encoding unit, etc., and can also be not limited to the texture information of the front and rear frame images, and can refer to the residual information between the encoding unit matching blocks, the median absolute deviation (MAD) and the like, wherein the MAD is used to represent the difficulty of residual encoding of the encoding block.

[0051] It should be noted that the inter-frame information reflects the inter-frame prediction correlation, so as to better evaluate the video encoding process.

[0052] It can be understood that the global code rate reference data can also include more types and more extensive data, and the above examples of the global code rate reference data are only used to illustrate the principle characteristics, but should not be understood as any limitation on its structure, and a person skilled in the art can select relevant types of global code rate reference data according to specific application scenarios and input them into the graph neural network individually or in combination, for example, the optimization setting for ROI encoding can be selected according to specific scenarios to improve the encoding effect, and since there is no mandatory dependency relationship between the encoding standard and the global code rate reference data, the specificity between the encoding standard and the global code rate reference data can not be considered, and the application scenarios are more extensive.

[0053] In one embodiment, the division of coding units in the to-be-compressed video can be determined based on the division information of the coding units. In one scenario, after the division of coding units is determined, the quantization parameters of each coding block in the coding unit are further determined, which is conducive to further determining the rate control parameters at the macroblock level. It can be understood that whether one of the two is confirmed separately or both are confirmed simultaneously, it will not affect the execution of the steps of this embodiment, but the corresponding emphasis is different, that is, the emphasis may be placed on controlling the division information of the coding units or the quantization parameters of each coding block, which is not limited in this embodiment. In addition, the rate-related data can also be data on the impact of the coding frame on the subjective compression quality. Although the rate control in the above embodiments is a strong constraint, during the iterative optimization of the rate parameter output, the impact of the coding frame on the subjective quality can be added as an input parameter to the network training process to solve the problem of excessive compression of the coding blocks in the NROI resulting in extreme deterioration of the subjective quality.

[0054] In one embodiment, step S100 can be presented as a specific function in a logical entity, which can be a separate physical device entity or a software entity on a host. The logical entity can be named a data preparation unit, which is to input the acquired global encoding reference data of the video to be compressed into the graph neural network, thereby obtaining the bit rate correlation data output by the graph neural network.

[0055] Step S200: determining a current bit rate parameter for controlling the video encoding bit rate according to the bit rate association data.

[0056] In one embodiment, a graph neural network is used to perform target constraint output on the global coding reference data of the video to be compressed, and bit rate correlation data under global optimization is obtained, which can reduce the impact of bit rate fluctuations caused by global errors. Then, based on the division information of the coding units and / or the quantization parameters of each coding block in the bit rate correlation data, the current bit rate parameters at the macroblock level suitable for the application scenario of the video to be compressed are obtained, which is beneficial to improving video coding efficiency and optimizing user viewing experience. It does not display the introduction of standard correlation information and can better adapt to multiple coding standards.

[0057] It can be understood that steps S100 and S200 have the following significant advantages:

[0058] Compared with the related art, this embodiment estimates the quantization parameter information under specific bit rate requirements by mathematically modeling the relationship between the statistical information of the coding block and the bit rate requirement. This embodiment does not require a unified coding standard and is suitable for video coding schemes with hybrid coding strategies, such as H.26x, VP9, ​​AV1, AVSx and other coding standards. There is no strong coupling relationship between the coding standard and the encoder capability, which makes it easier to implement hardware coding chip integration.

[0059] Compared with the related art, in which the target number of coding bits saved by NROI is allocated to the bit-coded ROI macroblock by superimposing ROI information, this embodiment takes a global perspective and fully considers the impact of visual overshoot, which can alleviate the excessive blur often caused by NROI and optimize the user video experience.

[0060] Compared with related technologies, the deep learning-based compression method realizes end-to-end encoding, such as the output of video compression parameters, which usually inputs the video output stream, or the estimation of network parameters, such as the use of statistical data such as confidence to evaluate the minimum bit rate. This embodiment can provide macroblock-level encoding parameters and does not need to rely on existing bit rate control methods. It can improve the adaptability to the scene and optimize the user video experience.

[0061] exist Figure 2 In the example of , when the rate-associated data includes division information of the coding unit and quantization parameters of each coding block, step S200 includes but is not limited to step S210.

[0062] In step S210, when the division information of the coding unit is determined, the quantization parameters of each coding block are trained based on the graph neural network to obtain the current bit rate parameters for controlling the video encoding bit rate.

[0063] In one embodiment, considering the scenario of determining the division information of the coding unit, the control adjustment of the specific bit rate is achieved by optimizing the quantization parameter configuration. In this case, the coding unit does not participate in the training process of the graph neural network as a fixed value, but the adjustment of the quantization parameter is achieved through the joint action of several other related global coding reference data. This adjustment method is highly targeted, and only the quantization parameter needs to be adjusted to achieve the output of the corresponding macroblock-level coding parameter, which is conducive to obtaining the current bit rate parameter more accurately and reasonably.

[0064] It can be understood that this embodiment takes into account the scenario where the coding unit is trainable. If conditions are sufficient, the results obtained based on advanced coding search can be used as true values ​​to participate in the training of the coding unit. This is not limited in this embodiment.

[0065] The above embodiment is described below with reference to specific examples.

[0066] Example 1:

[0067] like Figure 3 As shown, Figure 3 It is a schematic diagram of the structure of a graph neural network provided by one embodiment of the present invention.

[0068] exist Figure 3In the example, the graph neural network can be applied to, but not limited to, products or application equipment involving video encoding and decoding, such as terminals and smart interconnections. The global encoding reference data obtained and input this time includes bit rate constraint information, ROI information, reference frame information, current frame information, and texture statistics of the corresponding frame. Figure 3 The graph neural network shown applies texture statistical information based on the input global coding reference data, determines the division information of the coding unit and the quantization parameters of each coding block, and then trains the division information of the coding unit and the quantization parameters of each coding block by the graph neural network to output the required current bit rate parameters.

[0069] exist Figure 4 In the example, when the graph neural network is trained based on the acquired global encoding reference data, step S100 includes but is not limited to steps S110 to S120.

[0070] Step S110: Obtaining encoding frame information and historical bit rate parameters of the video to be compressed based on the graph neural network;

[0071] Step S120: Input the global coding reference data, coding frame information and historical bit rate parameters into the graph neural network, and output bit rate associated data. The historical bit rate parameters are the current bit rate parameters determined last time.

[0072] It should be noted that the graph neural network can be trained and constructed based on the acquired global coding reference data. After the training is completed, the global coding reference data is input into the constructed graph neural network. The constructed graph neural network can match the video coding requirements.

[0073] In one embodiment, consideration is given to optimizing input data based on global coding reference data, i.e., obtaining coding frame information and historical bitrate parameters of the video to be compressed through a graph neural network, and inputting the mixed global coding reference data into the graph neural network to obtain bitrate correlation data with better coding correlation; wherein, the coding frame information reflects the specific impact of the coding frame on the coding, and optimization based on the current bitrate parameters determined last time can take into account the historical determination scenario of the bitrate parameters, which is equivalent to further outputting bitrate correlation data based on the historical determination scenario of the bitrate parameters, thereby realizing optimized output of the bitrate parameters.

[0074] exist Figure 5 In the example, step S120 includes but is not limited to steps S121 to S122.

[0075] Step S121, determining encoding quality evaluation parameters according to encoding frame information and historical bit rate parameters;

[0076] Step S122, input the global code rate reference data and the coding quality evaluation parameter into the graph neural network, and output the code rate correlation data.

[0077] In an embodiment, the coding quality evaluation parameter is determined by the encoding frame information and the historical code rate parameter, and then the influence of the coding quality evaluation parameter is further matched with the influence of the global code rate reference data, so as to realize the optimized output of the code rate correlation data. It can be understood that when the code rate correlation data needs to be optimized, the coding quality evaluation parameter of the embodiment can be used as a new factor to realize the influence. In other words, if the code rate correlation data does not need to be further optimized, the coding quality evaluation parameter can be set to null. This is not limited in the embodiment.

[0078] It should be noted that in different application scenarios, the coding quality evaluation parameter determined is different because the obtained encoding frame information and historical code rate parameter are different. In addition, even in the same application scenario, different calculation methods can be used to obtain corresponding coding quality evaluation parameters, so as to output and optimize certain aspect or multiple aspects of the code rate correlation data according to the specific coding quality evaluation parameter. That is, each coding quality evaluation parameter can be different. This is not limited in the embodiment, and specific embodiments are given below for illustration.

[0079] In the example of Figure 6 , in the case that the encoding frame information includes the reference frame information and the coding quality evaluation parameter includes the first coding quality evaluation parameter, step S121 includes but is not limited to steps S1211 to S1213.

[0080] Step S1211, determining the encoding bitstream corresponding to the historical code rate parameter according to the historical code rate parameter;

[0081] Step S1212, decoding the encoding bitstream according to the reference frame information to obtain the reconstructed frame;

[0082] Step S1213, determining the first coding quality evaluation parameter according to the reconstructed frame.

[0083] In an embodiment, the encoding bitstream in the historical scene is determined and decoded to recover the reconstructed frame, so as to realize the reconstruction strategy targeting the recovered original frame. Since the reconstructed frame is associated with the reference frame information and the encoding bitstream corresponding to the historical code rate parameter, the reconstructed frame can represent the encoding condition of the historical scene and the encoding condition corresponding to the reference frame information. Under this condition, the first coding quality evaluation parameter determined based on the reconstructed frame has good forward propagation characteristics, which can meet the optimization training requirements based on the graph neural network, and is beneficial to improve the code rate parameter result output.

[0084] In the example of Figure 7In the example of FIG. 12, when there are multiple reconstructed frames, and each reconstructed frame corresponds to an encoded code stream, step S1213 includes but is not limited to steps S12131-S12132.

[0085] In step S12131, for each reconstructed frame, an encoded quality evaluation index corresponding to the reconstructed frame is obtained according to the reconstructed frame.

[0086] In step S12132, from the encoded quality evaluation indexes, a maximum encoded quality evaluation index is determined as the first encoded quality evaluation parameter.

[0087] In an embodiment, for each encoded code stream, the quality of the decoded data frame corresponding to the encoded code stream needs to be evaluated, that is, the encoded quality evaluation index corresponding to each reconstructed frame needs to be obtained, so that multiple encoded quality evaluation indexes can be obtained, and then the quality of the reconstructed frame in the current network environment is taken as a target function to update the training parameters of the graph neural network, the maximum encoded quality evaluation index is determined as the first encoded quality evaluation parameter, which indicates that the quality of the decoded data frame corresponding to the first encoded quality evaluation parameter is the largest, and therefore the graph neural network can be trained based on the parameter to optimize the code rate parameter output.

[0088] In Figure 8 the example of FIG. 12, step S12131 includes but is not limited to steps S12133-S12134.

[0089] In step S12133, a reconstructed quality parameter, a network lag parameter, and a switching condition parameter corresponding to the reconstructed frame are determined according to the reconstructed frame.

[0090] In step S12134, the reconstructed quality parameter, the network lag parameter, and the switching condition parameter are weighted and superimposed to obtain an encoded quality evaluation index corresponding to the reconstructed frame.

[0091] In an embodiment, by introducing the weighted and superimposed values of the reconstructed quality parameter, the network lag parameter, and the switching condition parameter, the encoded quality evaluation index corresponding to the reconstructed frame can be accurately obtained, and the encoded quality evaluation index is only related to the quality parameter content of the reconstructed frame itself and does not involve other impurities for calculation, so the error fluctuation is relatively small.

[0092] The following specific examples are given to illustrate the principles of the embodiments.

[0093] Example II:

[0094] As Figure 9 shown, Figure 9 is an execution flowchart for determining the first encoded quality evaluation parameter provided by an embodiment of the present application.

[0095] In Figure 9In the example, perform the following steps in sequence:

[0096] Step S300: Obtaining an encoded bitstream corresponding to the historical bitrate parameters obtained from the graph neural network;

[0097] Step S400: referencing a reference frame, decoding the encoded bitstream through a decoder to generate a decoding result, and obtaining a reconstructed frame;

[0098] Step S500: Determine a first encoding quality assessment parameter based on the reconstructed frame.

[0099] Among them, corresponding to step 3, the quality of the restored frame in the current network environment is used as the objective function to update the network parameters. For example, the weighted combination of reconstruction quality, network jamming parameters and switching status can be comprehensively introduced as the overall quality of experience (QoE) evaluation index, that is,

[0100]

[0101] R(n) can use non-reference image quality evaluation indicators, including but not limited to Information Fidelity Criterion (IFC), Deep CNN-Based Blind Image Quality Predictor (DIQA), etc.

[0102] It is understandable that the subjective quality of the reconstructed frames can also be evaluated by the Generative Adversarial Network (GAN), and reference can be made to high-quality reconstruction network architectures such as the Enhanced Generative Adversarial Network (ESRGAN).

[0103] The coding strategy proposed in this embodiment requires that the video to be compressed be coded by region. Different coding parameters and strategies are designed based on the differences in regional information (such as ROI, texture statistics, etc.). The quality degradation of the final output coded frame is minimized under the overall bit rate control. The following objective function is considered:

[0104]

[0105] Among them, GNN′(X) represents the coded stream output by this example, and there are multiple coded streams. Q(GNN′(X)) represents the quality of the data frame obtained by decoding the coded stream. The constraint condition is BD GNN′(X)≤RATE, the bitrate should not exceed the specified target bitrate. Each encoding solution that satisfies the constraint is considered an action, and the discriminant function f is used as the evaluation mechanism. The goal is to find the largest f. Under this model, a graph neural network can be trained using reinforcement learning to maximize video quality under specific bitrate requirements.

[0106] exist Figure 10 In the example, when the coded frame information further includes current frame information and the coding quality assessment parameter further includes a second coding quality assessment parameter, step S121 further includes but is not limited to step S1214.

[0107] Step S1214 : performing differential processing on the reconstructed frame information and the current frame information to obtain a second encoding quality assessment parameter, wherein the reconstructed frame information corresponds to the reconstructed frame.

[0108] In one embodiment, after determining the reconstructed frame, the obtained reconstructed frame information is differentiated in conjunction with the current frame information, so as to take into account the encoding situation corresponding to the current frame information, and obtain a second encoding quality evaluation parameter that meets the requirements, which can meet the optimization training requirements based on the graph neural network and is conducive to improving the bit rate parameter result output. Among them, the objective function can be obtained based on the differential processing, and then the encoding result is evaluated based on the determined objective function. The following is a specific example to illustrate the principle of this embodiment.

[0109] Example 3:

[0110] like Figure 11 As shown, Figure 11 This is an execution flow chart of determining a second encoding quality assessment parameter provided by an embodiment of the present invention.

[0111] exist Figure 11 In the example, perform the following steps in sequence:

[0112] Step S600: Obtaining an encoded bitstream corresponding to the historical bitrate parameters obtained from the graph neural network;

[0113] Step S700: referencing a reference frame, decoding the encoded bitstream through a decoder to generate a decoding result;

[0114] Step S800: Compare the decoding result with the true value of the current frame, calculate the difference cost f, and obtain Loss (ie, the second encoding quality evaluation parameter).

[0115] The method for obtaining the Loss is determined according to the specific application scenario, which is not limited in this embodiment and is described below with examples.

[0116] f=||x'-x||1

[0117] As shown in the above formula, the L1 norm of the reconstructed image x' and the uncompressed image x is used as the Loss, or an implicit discrimination method can also be used, for example, a discrimination network based on the idea of GAN is designed to analyze the quality of the encoded image, that is,

[0118] f = ||g(h'(h(x)))-g(x) || 1

[0119] Where h and h' represent the encoding unit and the decoding unit respectively, since h is lossy compression, the recovered image quality is degraded, and by reconstructing the objective function g(x), that is, the discriminator part of GAN, or the output of the discriminator network of ESRGAN can also be used to evaluate the encoding result, so as to realize the maximum preservation of video quality under the requirement of a certain code rate.

[0120] It can be understood that the Loss calculation based on the current frame and the reconstructed frame can also use a variety of similar schemes, for example, in step 3, the L2 norm of the reconstructed image x' and the uncompressed image x is used as the Loss, etc.

[0121] It should be noted that the execution process of examples two and three can be presented in a logical entity as a specific function, which can be a separate physical device entity, or a software entity on a host, and the logical entity can be named a model training unit, which determines the first encoding quality evaluation parameter according to the reconstructed frame, and differentially processes the reconstructed frame information and the current frame information to obtain the second encoding quality evaluation parameter.

[0122] In the example of Figure 12 , when the code rate related data includes the second code rate related data, step S122 includes but is not limited to step S1221.

[0123] Step S1221 inputs the global code rate reference data and the second encoding quality evaluation parameter into the graph neural network to obtain the second code rate related data.

[0124] In an embodiment, by inputting the global code rate reference data and the second encoding quality evaluation parameter into the graph neural network, the second code rate related data corresponding to the second encoding quality evaluation parameter is obtained, and compared with the original code rate related data, using the second encoding quality evaluation parameter as a training parameter to optimize the graph neural network can obtain second code rate related data with better optimization effect, which is conducive to improving the video compression effect.

[0125] In the example of Figure 13 , when the code rate related data includes the first code rate related data, step S122 includes but is not limited to step S1222.

[0126] Step S1222: For each encoded bitstream, the global bitrate reference data and the first encoding quality assessment parameter are input into the graph neural network to obtain first bitrate associated data corresponding to the encoded bitstream.

[0127] In one embodiment, for each encoded bitstream, the global bitrate reference data and the first encoding quality assessment parameter are input into the graph neural network to obtain the first bitrate associated data corresponding to each encoded bitstream. That is, in a specific application scenario, during a video compression process, the first bitrate associated data corresponding to each encoded bitstream can be controlled and adjusted separately, which can avoid homogenization and significantly improve the video compression effect.

[0128] exist Figure 14 In the example, step S100 also includes but is not limited to step S900.

[0129] Step S900: Upon receiving resource limitation information corresponding to the graph neural network, scale compression processing is performed on the graph neural network.

[0130] In one embodiment, resource limitation information may be formed in a scenario where resource limitations are imposed on the application platform. In this scenario, the graph neural network is subjected to scale compression processing according to the requirements of the application scenario, including but not limited to distillation, quantization, pruning, and dynamic network design, etc., to reduce the scale and computing power requirements of the overall graph neural network model. Accordingly, in the scenario of resource expansion processing, the graph neural network can be scaled up, or a new graph neural network that meets the requirements can be used to replace the original graph neural network.

[0131] exist Figure 15 In the example, step S200 also includes but is not limited to step S1000.

[0132] Step S1000: upon receiving model adaptation information corresponding to the current bit rate parameter, optimizing the current bit rate parameter according to the model adaptation information.

[0133] In one embodiment, the model adaptation information may be formed in a scenario where the network transmission environment is constrained. In this scenario, the current bitrate parameters are optimized based on the model adaptation information, including but not limited to optimizing encoding parameters, considering reducing the bitrate at the expense of subjective quality, etc., so as to adapt the network structure and model parameters.

[0134] It should be noted that the step S900 and the step S1000 can be presented in a logical entity as a specific function, which can be a separate physical device entity, or a software entity on a host, and the logical entity can be named as an inference application unit, which is deployed in combination with a lightweight strategy, performs scale compression processing on the graph neural network in a resource-limited scenario, optimizes the current code rate parameter in a scenario where the network transmission environment is constrained, and achieves the purpose of reducing the power consumption of the model.

[0135] In addition, with reference to Figure 16 , one embodiment of the present application also provides a video code rate control device 100, which comprises a memory 110, a processor 120 and a computer program stored in the memory 110 and executable on the processor 120.

[0136] The processor 120 and the memory 110 can be connected through a bus or other means.

[0137] The non-transitory software program and instructions required to implement the video code rate control method of the above-mentioned embodiments are stored in the memory 110, and when executed by the processor 120, the video code rate control method of each embodiment is executed, for example, the method steps S100 to S200 in the above-mentioned Figure 1 , the method step S210 in the above-mentioned Figure 2 , the method steps S110 to S120 in the above-mentioned Figure 4 , the method steps S121 to S122 in the above-mentioned Figure 5 , the method steps S1211 to S1213 in the above-mentioned Figure 6 , the method steps S12131 to S12132 in the above-mentioned Figure 7 , the method steps S12133 to S12134 in the above-mentioned Figure 8 , the method steps S300 to S500 in the above-mentioned Figure 9 , the method step S1214 in the above-mentioned Figure 10 , the method steps S600 to S800 in the above-mentioned Figure 11 , the method step S1221 in the above-mentioned Figure 12 , the method step S1222 in the above-mentioned Figure 13 , the method step S900 in the above-mentioned Figure 14 , or the method step S1000 in the above-mentioned Figure 15 .

[0138] The device embodiments described above are only schematic, and the units described as separate components can or can not be physically separate, that is, they can be located in one place, or distributed on multiple network units. Part or all of the modules can be selected according to actual needs to achieve the purpose of the present embodiment scheme.

[0139] Further, an embodiment of the present application also provides a computer readable storage medium storing computer executable instructions, which, when executed by a processor 120 or a controller, such as a processor 120 in the above device embodiment, can cause the processor 120 to perform the video rate control method in the above embodiments, such as performing the method steps S100-S200 in the method of Figure 1 Figure 2 the method step S210 in the method of Figure 4 the method steps S110-S120 in the method of Figure 5 the method steps S121-S122 in the method of Figure 6 the method steps S1211-S1213 in the method of Figure 7 the method steps S12131-S12132 in the method of Figure 8 the method steps S12133-S12134 in the method of Figure 9 the method steps S300-S500 in the method of Figure 10 the method step S1214 in the method of Figure 11 the method steps S600-S800 in the method of Figure 12 the method step S1221 in the method of Figure 13 the method step S1222 in the method of Figure 14 the method step S900 in the method of Figure 15 the method step S1000 in the method of

[0140] ​Those skilled in the art will appreciate that all or some of the steps and systems in the method disclosed above can be implemented as software, firmware, hardware, and appropriate combinations thereof. Some physical components or all physical components can be implemented as software executed by a processor, such as a central processing unit, a digital signal processor, or a microprocessor, or implemented as hardware, or implemented as an integrated circuit, such as an application-specific integrated circuit. Such software can be distributed on a computer-readable medium, and the computer-readable medium can include computer storage media (or non-transitory media) and communication media (or temporary media). As known to those skilled in the art, the term computer storage media is included in any method or technology for storing information (such as computer-readable instructions, data structures, program modules, or other data) and is volatile and non-volatile, removable, and non-removable. Computer storage media includes, but is not limited to, RAM, ROM, EEPROM, flash memory, or other memory technology, CD-ROM, digital versatile disks (DVD), or other optical disk storage, magnetic cassettes, magnetic tapes, disk storage, or other magnetic storage devices, or any other medium that can be used to store desired information and can be accessed by a computer. Furthermore, as is well known to those skilled in the art, communication media typically includes computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transport mechanism, and may include any information delivery media.

[0141] The above is a specific description of the preferred embodiments of the present invention, but the present invention is not limited to the above embodiments. Those skilled in the art can make various equivalent modifications or substitutions without violating the spirit of the present invention. These equivalent modifications or substitutions are all included in the scope defined by the claims of the present invention.

Claims

1. A video bit rate control method, comprising: Input the obtained global encoding reference data of the video to be compressed into the graph neural network and output the bit rate correlation data; Determining a current bit rate parameter for controlling a video encoding bit rate according to the bit rate associated data; The global coding reference data is used to characterize the compression quality of the video to be compressed, and the bit rate associated data includes at least one of the following types: division information of the coding unit, and quantization parameters of each coding block in the coding unit.

2. The rate control method according to claim 1, wherein: The graph neural network is trained based on the acquired global coding reference data; the acquired global coding reference data of the video to be compressed is input into the graph neural network, and the bit rate associated data is output, including: Obtaining encoding frame information and historical bit rate parameters of the video to be compressed based on the graph neural network; The global coding reference data, the coding frame information and the historical bit rate parameters are input into the graph neural network, and bit rate association data is output, where the historical bit rate parameters are the current bit rate parameters determined last time.

3. The rate control method according to claim 1, wherein: The bit rate associated data includes division information of the coding unit and quantization parameters of each coding block; Determining a current bit rate parameter for controlling the video encoding bit rate according to the bit rate associated data includes: Based on the quantization parameters of the respective coding blocks and the division information of the coding units, a current bit rate parameter for controlling the video encoding bit rate is obtained.

4. The rate control method according to claim 1, wherein: The global coding reference data includes at least one of the following types: Rate constraint information associated with the coding standard; Region of interest ROI information; Encoding type information; Encoder information; Encoding frame constraint information; Coded frame statistics; Inter-frame information.

5. The rate control method according to claim 2, wherein: The step of inputting the global coding reference data, the coding frame information, and the historical bit rate parameter into the graph neural network and outputting bit rate associated data comprises: Determining a coding quality assessment parameter according to the coding frame information and the historical bit rate parameter; The global coding reference data and the coding quality assessment parameters are input into the graph neural network, and bit rate associated data is output.

6. The rate control method according to claim 5, wherein: The coded frame information includes reference frame information, and the coding quality assessment parameter includes a first coding quality assessment parameter; and determining the coding quality assessment parameter according to the coded frame information and the historical bit rate parameter includes: Determining, according to the historical bit rate parameters, an encoding stream corresponding to the historical bit rate parameters; Decoding the coded code stream according to the reference frame information to obtain a reconstructed frame; The first encoding quality assessment parameter is determined according to the reconstructed frame.

7. The rate control method according to claim 6, wherein: The coded frame information further includes current frame information, and the coding quality assessment parameter further includes a second coding quality assessment parameter; and determining the coding quality assessment parameter according to the coded frame information and the historical bit rate parameter further includes: Perform differential processing on the reconstructed frame information and the current frame information to obtain the second encoding quality assessment parameter, wherein the reconstructed frame information corresponds to the reconstructed frame.

8. The rate control method according to claim 6, wherein: There are multiple reconstructed frames, each of which corresponds to one encoded code stream; and determining the first encoding quality assessment parameter according to the reconstructed frames includes: For each of the reconstructed frames, obtaining a coding quality assessment index corresponding to the reconstructed frame according to the reconstructed frame; From the various encoding quality assessment indicators, the largest encoding quality assessment indicator is determined as the first encoding quality assessment parameter.

9. The rate control method according to claim 8, wherein: The obtaining, according to the reconstructed frame, a coding quality assessment indicator corresponding to the reconstructed frame includes: Determining, according to the reconstructed frame, a reconstruction quality parameter, a network stall parameter, and a switching status parameter corresponding to the reconstructed frame; The reconstruction quality parameter, the network stall parameter, and the switching status parameter are weightedly superimposed to obtain a coding quality evaluation index corresponding to the reconstructed frame.

10. The rate control method according to claim 8, wherein: The bit rate associated data includes first bit rate associated data; inputting the global coding reference data and the coding quality assessment parameter into the graph neural network and outputting the bit rate associated data includes: For each of the encoded code streams, the global encoding reference data and the first encoding quality assessment parameter are input into the graph neural network to obtain the first bit rate associated data corresponding to the encoded code stream.

11. The rate control method according to claim 7, wherein: The bit rate associated data includes second bit rate associated data; the inputting the global coding reference data and the coding quality assessment parameter into the graph neural network and outputting the bit rate associated data includes: The global coding reference data and the second coding quality assessment parameter are input into the graph neural network to obtain the second bit rate associated data.

12. The rate control method according to claim 1, wherein: The method further includes: inputting the acquired global coding reference data of the video to be compressed into the graph neural network and outputting the bit rate associated data before outputting the bit rate associated data; Upon receiving resource limitation information corresponding to the graph neural network, the graph neural network is subjected to scale compression processing.

13. The rate control method according to claim 1, wherein: After determining the current bit rate parameter for controlling the video encoding bit rate according to the bit rate associated data, the method further includes: When the model adaptation information corresponding to the current bit rate parameter is received, the current bit rate parameter is optimized according to the model adaptation information.

14. A video bit rate control device, comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the video bit rate control method according to any one of claims 1 to 13 when executing the computer program.

15. A computer-readable storage medium storing computer-executable instructions, wherein the computer-executable instructions are used to execute the video bit rate control method according to any one of claims 1 to 13.

Citation Information

Patent Citations

  • Video code rate control method and video coding device

    CN106331704A

  • Video coding intra-frame code rate control method based on deep reinforcement learning

    CN111294595A