Video encoding method, video encoder, electronic equipment and storage medium
By dividing the super block into multiple blocks and adjusting the initial indicators according to the complexity and reference degree of block content, the encoding quality and inefficiency caused by using the same QP in the super block is solved, and more efficient video encoding is achieved.
Patent Information
- Application Number
- CN202411978347.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-30
- Publication Date
- 2025-05-06
AI Technical Summary
In the prior art, superblocks (SBs) have both texture flat areas and texture uneven areas, and the use of the same quantization parameter (QP) leads to low encoding quality and encoding efficiency.
The process of fine-tuning the quantization parameter (QP) is obtained by dividing the superblock in the video frame to be encoded into multiple blocks of specified pixel sizes, and adjusting the initial indicators using the complexity of the block content and the degree of reference by other video frames.
It improves encoding quality and coding efficiency, and solves the problem of encoding quality and inefficiency caused by the use of the same QP in different regions in SB.
Smart Images

Figure CN119946285A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of video technology, and in particular to a video encoding method, a video encoder, an electronic device and a storage medium. Background Art
[0002] Video encoding is a technical process that compresses raw video data (i.e., usually uncompressed, containing a large amount of visual information, with a very large amount of data, requiring a lot of resources to process and store. For example, a high-resolution video may contain hundreds of megabytes of data per second) for storage or transmission. Through video encoding, the raw video data can be compressed into a smaller file size while maintaining the video quality as much as possible. The application architecture of video coding technology can be as follows: Figure 1 As shown, the terminal 10, the terminal 20, and the server 30 are included. The terminal 10 uploads the encoded video to the server, and the terminal 20 downloads the encoded video from the server 30 and displays it. However, currently, the encoder mainly implements adaptive quantization based on blocks (for example, blocks of 8*8 pixels). However, with the introduction of super blocks (SB), this block-based adaptive quantization method is no longer applicable. For this technical problem, the relevant technology has not yet proposed an effective solution. Summary of the invention
[0003] Embodiments of the present application provide a video encoding method, a video encoder, an electronic device, and a storage medium to solve one or more of the above-mentioned technical problems.
[0004] In a first aspect, an embodiment of the present application provides a video encoding method, comprising: for a super block SB in a video frame to be encoded, dividing a target area in the SB into multiple blocks of specified pixel sizes; for any one of the multiple blocks, adjusting the initial index of the block by a target adjustment factor corresponding to the block, to obtain target indexes corresponding to the multiple blocks respectively, wherein the target adjustment factor includes at least a first adjustment factor and a second adjustment factor, the first adjustment factor is used to characterize the complexity of the block content, and the second adjustment factor is used to characterize the degree to which the block is referenced by other video frames in the video frame to be encoded, and there is a preset relationship between the value of the initial index and the value of a block-level quantization parameter QP; determining the target index of the target area by the target indexes corresponding to the multiple blocks respectively, and encoding the target area based on the target index of the target area.
[0005] In a second aspect, an embodiment of the present application provides a video encoder, comprising at least a pre-processing module, wherein the pre-processing module performs encoding using the above-mentioned video encoding method.
[0006] In a third aspect, an embodiment of the present application provides an electronic device, including a memory, a processor, and a computer program stored in the memory, wherein the processor implements any of the above methods when executing the computer program.
[0007] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method described in any one of the above is implemented.
[0008] In a fifth aspect, an embodiment of the present application provides a computer program product, including computer instructions, which implement any of the methods described above when executed by a processor.
[0009] Compared with the related art, this application has the following advantages:
[0010] According to an embodiment of the present application, for a super block SB in a video frame to be encoded, a target area in the SB is divided into multiple blocks of specified pixel sizes; for any one of the multiple blocks, an initial indicator of the block is adjusted by a target adjustment factor corresponding to the block to obtain target indicators respectively corresponding to the multiple blocks, wherein the target adjustment factor includes at least a first adjustment factor and a second adjustment factor, the first adjustment factor is used to characterize the complexity of the content of the block, and the second adjustment factor is used to characterize the degree to which the block is referenced by other video frames in the video frame to be encoded, and there is a preset relationship between the value of the initial indicator and the value of the block-level quantization parameter QP; the target indicator of the target area is determined by the target indicators respectively corresponding to the multiple blocks, and the target area is encoded based on the target indicator of the target area. That is to say, the embodiment of the present application achieves the purpose of fine-tuning QP by jointly adjusting the initial indicators of different areas in SB through a first adjustment factor used to characterize the complexity of block content and a second adjustment factor used to characterize the degree to which the block is referenced by other video frames in the video frame to be encoded, thereby solving the technical problem in the related art that the SB has both flat texture areas and uneven texture areas, and the use of the same QP leads to low encoding quality and encoding efficiency, thereby achieving the technical effect of improving encoding quality and encoding efficiency.
[0011] The above description is only an overview of the technical solution of the present application. In order to more clearly understand the technical means of the present application, it can be implemented in accordance with the contents of the specification. In order to make the above and other purposes, features and advantages of the present application more obvious and easy to understand, the specific implementation methods of the present application are listed below. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] In the accompanying drawings, unless otherwise specified, the same reference numerals throughout the multiple drawings represent the same or similar parts or elements. These drawings are not necessarily drawn to scale. It should be understood that these drawings only depict some embodiments according to the present application and should not be regarded as limiting the scope of the present application.
[0013] Figure 1 A schematic diagram of a video encoding architecture in a related technology provided in an embodiment of the present application is shown;
[0014] Figure 2 A schematic diagram of a video encoding method provided in an embodiment of the present application is shown;
[0015] Figure 3 A structural block diagram of a video encoding device provided in an embodiment of the present application is shown;
[0016] Figure 4 A structural block diagram of a video encoder provided in an embodiment of the present application is shown; and
[0017] Figure 5 A block diagram of an electronic device used to implement an embodiment of the present application is shown. DETAILED DESCRIPTION
[0018] In the following, only some exemplary embodiments are briefly described. As those skilled in the art will appreciate, the described embodiments may be modified in various ways without departing from the concept or scope of the present application. Therefore, the drawings and descriptions are considered to be exemplary in nature and not restrictive.
[0019] To facilitate understanding of the technical solutions of the embodiments of the present application, the following describes the related technologies of the embodiments of the present application. The following related technologies can be combined with the technical solutions of the embodiments of the present application as optional solutions, and they all belong to the protection scope of the embodiments of the present application.
[0020] Explanation of terms
[0021] SB is a larger block structure that can better adapt to the diversity of different image regions compared to the traditional smaller block (8x8 pixel) structure. In the AV1 encoder, super blocks are usually 128x128 pixels or 64x64 pixels in size. The introduction of super blocks is to more effectively handle complex scenes such as high-speed motion and static backgrounds, reduce encoding complexity and improve compression efficiency.
[0022] AV1 encoder is an advanced video compression technology developed by the Alliance for Open Media (AOMedia). The encoder is designed to improve video compression efficiency, provide higher quality video content, while reducing bandwidth usage.
[0023] QP is a parameter that controls the quantization level in the video encoding process. In video encoding, quantization is the process of mapping transform coefficients from the continuous domain to the discrete domain. The QP value determines the fineness of this mapping. The finer the quantization process (i.e., the lower the QP value), the more image details are retained, but the compression rate is lower. Conversely, the coarser the quantization process (i.e., the higher the QP value), the more image details are lost, but the compression rate is higher.
[0024] Delta Quantization Parameter (dQP) is a supplementary adjustment relative to a certain baseline QP. It enables the encoder to more flexibly refine the quantization level during the encoding process and perform more precise quantization control on different image regions or blocks. This refined control brings higher encoding efficiency and more optimized visual quality.
[0025] Adaptive quantization (AQ) is based on the fact that the human eye is insensitive to distortion in areas with complex textures. It increases the QP value for areas with complex textures and decreases the QP value for flat areas. The algorithm represents the spatial texture complexity of the video by calculating the variance or edge density.
[0026] Structural Similarity Index (SSIM) is an indicator used to measure the similarity between two images, with a particular focus on the perceived quality of the human eye. Unlike the traditional Peak Signal-to-Noise Ratio (PSNR), which only evaluates quality through pixel-to-pixel differences, SSIM measures similarity through a comprehensive evaluation of image brightness, contrast, and structure to ensure that it is more in line with the subjective perception of the human eye. The SSIM value ranges from 0 to 1, with 1 indicating that the two images are exactly the same and 0 indicating that they are completely different. The AV1 encoder, as an advanced compression technology, often uses the SSIM indicator to measure its encoding effect. During the compression process, the AV1 encoder continuously evaluates the SSIM value between the original video and the compressed video to ensure that the visual quality after compression meets the expected standard. A higher SSIM value indicates that the compressed video quality is better. The AV1 encoder can dynamically adjust the encoding strategy based on the SSIM value, such as selecting different quantization parameters, prediction modes, and intra / inter coding methods to achieve the optimal balance between compression efficiency and video quality. When developing and using the AV1 encoder, the SSIM indicator can be used to tune the encoding parameters and find the best configuration for different application scenarios. For example, for high-quality streaming, higher encoding parameters may be set to obtain a higher SSIM value.
[0027] In a related technology before this application, AQ is implemented based on block variance. Specifically, based on the current block variance and the average variance of the entire frame, the difference between the current block and the average value of the entire frame is calculated, and then the difference is mapped to the dQP. However, there are both texture flat areas (i.e., parts of the image with small or more uniform changes) and texture uneven areas (i.e., parts of the image with large or more complex changes) in SB, and it is not reasonable to use the same QP.
[0028] In another related technology before the present application, for frames at high temporal levels (i.e., frames at a higher level in a hierarchical coding structure, used to improve temporal resolution, increase smoothness and detail presentation of the video. They are located at a higher level of the temporal hierarchy, frequently inserted between lower-level frames, and rely on lower-level frames for encoding), the QP is directly adjusted so that the transmitted bit stream is occupied by more dQP, while the bits occupied by the frames at the high temporal level themselves are very small, and the ultimate cost-effectiveness of this method is not high. In addition, if the quality of frames at low temporal levels (referring to frames at a lower level in hierarchical coding, especially in a temporal coding structure, which provide a stable basis for decoding and reconstruction of video content) is not high, then increasing the quality of frames at high temporal levels will not have a beneficial effect on improving coding quality and coding efficiency.
[0029] In view of this, the embodiments of the present application provide a video encoding method to solve all or part of the above-mentioned technical problems. The application scenarios of the method may be video on demand (Video on Demand, VOD) (allowing users to watch video content at any time according to their needs without being restricted by traditional TV program schedules. This service is provided through the Internet, and users can choose the movies, TV series, documentaries, variety shows, etc. they want to watch, and play them instantly on their devices (such as smart TVs, computers, tablets, mobile phones, etc.), video conferencing, online live broadcasts, security monitoring, virtual reality (Virtual Reality, VR) / augmented reality (Augmented Reality, AR), mobile video communications, telemedicine, cloud games, digital TV and broadcasting, education and training, etc. In other words, the video encoding method provided in the embodiments of the present application can be applied to related technologies such as video compression coding and storage. For example Figure 2 As shown, the above video encoding method may include:
[0030] S202: for a super block SB in a video frame to be encoded, divide a target area in the SB into a plurality of blocks of a specified pixel size.
[0031] The type of the video frame to be encoded may be a frame of a low temporal layer or a frame of a high temporal layer. The blocks of the above-mentioned multiple specified pixel sizes may be fixed-size blocks, for example, 16*16 pixels, 8*8 pixels, and the appropriate block size is selected according to the coding standard and application requirements. It may also be a variable-size block, and a block partitioning technique of size (such as quadtree partitioning) may be used to dynamically adjust the block size according to regional characteristics. Exemplarily, assuming that the target area is 64*64 pixels, the above-mentioned block is a fixed-size block, and is 16*16 pixels, then it can be divided into 16 blocks of 16*16 pixels. For another example, assuming that the target area is 64*64 pixels, and the above-mentioned block is a variable-size block, then it can be divided into 4 blocks of 32*32 pixels in the first layer, and the 32*32 pixel block is divided into 4 blocks of 16*16 pixels in the second layer, and the 16*16 pixel block is divided into 4 blocks of 8*8 pixels in the third layer.
[0032] The target region may be determined based on motion estimation, edge and texture complexity, etc. The number of the target region may include multiple regions, for example, it may be a region with flat texture or a region with uneven texture in the SB.
[0033] S204, for any block among the multiple blocks, adjust the initial indicator of the block by the target adjustment factor corresponding to the block to obtain the target indicators respectively corresponding to the multiple blocks, wherein the target adjustment factor includes at least a first adjustment factor and a second adjustment factor, the first adjustment factor is used to characterize the complexity of the content of the block, and the second adjustment factor is used to characterize the degree to which the block is referenced by other video frames in the video frame to be encoded, and there is a preset relationship between the value of the initial indicator and the value of the block-level quantization parameter QP.
[0034] Optionally, in an embodiment of the present application, in addition to the first adjustment factor and the second adjustment factor, the target adjustment factor may also include other adjustment factors that affect the initial indicator, which may be determined according to the application scenario and requirements, and are not limited here. The above-mentioned initial indicator is an indicator that can affect the QP setting, including but not limited to λ, wherein λ is a parameter related to rate-distortion optimization (RDO), reflecting the cost function in encoding, which is used to weigh bit rate and distortion in the video encoding process, thereby achieving a balance between encoding efficiency and video quality. Generally, the larger λ is, the more emphasis will be placed on compression efficiency rather than quality in the encoding process, and the smaller λ is, the more emphasis will be placed on quality. In the AV1 encoder, the relationship between λ and QP is usually represented by an empirical formula or curve. This formula can ensure that under different QPs, the selection of λ can most effectively optimize rate distortion. Exemplarily, in an embodiment of the present application, the relationship between λ and QP can be shown as Formula 1:
[0035]
[0036] Where α is a proportional constant specific to the encoder and is used to calibrate the starting value of λ. 0 and K 1 is a control parameter used to adjust the scale and speed of λ.
[0037] In the embodiment of the present application, the above-mentioned λ belongs to the video encoder end and does not need to be encoded into the bit stream and transmitted. Therefore, the embodiment of the present application adjusts λ to achieve fine-tuning of QP, and does not need to calculate and store dQP, thereby saving the bit overhead of dQP.
[0038] S206: Determine a target indicator of the target area through the target indicators respectively corresponding to the multiple blocks, and encode the target area based on the target indicator of the target area.
[0039] In an embodiment of the present application, the target indicator of the target area can be determined by the target indicators corresponding to the multiple blocks respectively, which can be: obtaining the geometric mean of the target indicators corresponding to the multiple blocks respectively, and determining the target indicator of the target area based on the geometric mean. It can also be: obtaining the maximum and minimum values of the target indicators corresponding to the multiple blocks respectively, and determining the target indicator of the target area based on the maximum and minimum values. It can also be: using a machine learning model (such as a regression model, a neural network, etc.) to predict the target indicator of the target area based on historical data and the target indicators corresponding to the input multiple blocks. In an embodiment of the present application, while maintaining the rationality of the calculation and maximizing the specific performance or image quality requirements, the target indicator of the target area can be determined by selecting a suitable method according to the specific application scenario and needs, and no specific limitation is made here.
[0040] Through the above steps S202 to S206, for the super block SB in the video frame to be encoded, the target area in the SB is divided into multiple blocks of specified pixel sizes; for any block in the multiple blocks, the initial index of the block is adjusted by the target adjustment factor corresponding to the block, and the target indexes corresponding to the multiple blocks are obtained, wherein the target adjustment factor includes at least a first adjustment factor and a second adjustment factor, the first adjustment factor is used to characterize the complexity of the content of the block, and the second adjustment factor is used to characterize the degree to which the block is referenced by other video frames in the video frame to be encoded, and there is a preset relationship between the value of the initial index and the value of the block-level quantization parameter QP; the target index of the target area is determined by the target indexes corresponding to the multiple blocks, and the target area is encoded based on the target index of the target area. That is to say, the embodiment of the present application achieves the purpose of fine-tuning QP by jointly adjusting the initial indicators of different areas in SB through a first adjustment factor used to characterize the complexity of block content and a second adjustment factor used to characterize the degree to which the block is referenced by other video frames in the video frame to be encoded, thereby solving the technical problem in the related art that the SB has both flat texture areas and uneven texture areas, and the use of the same QP leads to low encoding quality and encoding efficiency, thereby achieving the technical effect of improving encoding quality and encoding efficiency.
[0041] In a possible implementation, for any block among the multiple blocks, adjusting the initial index of the block by using the target adjustment factor corresponding to the block to obtain the target index of the block may include:
[0042] S11, determining a first adjustment factor corresponding to the block according to the variance corresponding to the block.
[0043] Among them, the variance var corresponding to the block can be determined by the following formula 2, where a i represents the value of the i-th pixel in the block, and N is the number of all pixels in the block:
[0044]
[0045] Optionally, in an embodiment of the present application, the method of determining the first adjustment factor corresponding to the block based on the variance corresponding to the block may be: mapping the variance and the first adjustment factor to a specified functional relationship, such as a linear function and a similar deformation of the linear function, etc. The specific method may be determined according to the application scenario and data characteristics, and is not limited here.
[0046] S12: Integrate the first adjustment factor corresponding to the block and the second adjustment factor corresponding to the block to obtain a target adjustment factor corresponding to the block.
[0047] Optionally, in the embodiment of the present application, the first adjustment factor and the second adjustment factor may be integrated by integrating the two adjustment factors through a custom function, which may be a linear or nonlinear function. The specific function may be determined according to the application scenario and data characteristics, and is not limited here.
[0048] S13, adjusting the initial index of the block by using the target adjustment factor corresponding to the block to obtain the target index corresponding to the block.
[0049] Optionally, in an embodiment of the present application, the method of adjusting the initial index of the block by the target adjustment factor corresponding to the block can be to integrate the target adjustment factor and the initial index by a custom function, which can be a linear or nonlinear function, which can be determined according to the application scenario and data characteristics. For example, assuming that the initial index is λ1 and the target adjustment factor is scale, then the target index λ2=scale*λ1.
[0050] In order to optimize the Structural Similarity Index (SSIM), the embodiment of the present application further proposes to determine the first adjustment factor corresponding to the block by the variance corresponding to the block, which may include: S111, establishing a mapping relationship between the first adjustment factor and the variance; S112, adding a first preset offset based on the mapping relationship to determine the first adjustment factor, wherein the first preset offset is determined based on the SSIM. That is, when the first adjustment factor is scale aq (the complexity of the block content determined by the spatial domain information), the variance is var, and the first preset offset is a constant C 1 When , the first adjustment factor can be determined by the following formula 3:
[0051] scale aq =2*var+C 1 (Formula 3)
[0052] Among them, constdoubleC 1 =(0.03*0.03*255*255)*(1<<2*(bitdepth-8)), where bitdepth is the bit width of the source video file, for example, it can be 10 bits or 8 bits.
[0053] In order to effectively integrate the influence of multiple factors, the embodiment of the present application proposes to integrate the first adjustment factor corresponding to the block and the second adjustment factor corresponding to the block to obtain the target adjustment factor corresponding to the block, which may include: S121, determining the product of the first adjustment factor and the second adjustment factor as the target adjustment factor. That is, when the first adjustment factor is scale aq, the second adjustment factor is scale tpl (Temporal Prediction List (TPL) adjustment factor, that is, the degree to which the current block is referenced by other video frames), the target adjustment factor scale can be determined by formula 4:
[0054] scale=scale aq *scale tpl (Formula 4)
[0055] Taking into account that the geometric mean is more effective in processing proportional data or data of relative size, it emphasizes the overall multiplication relationship rather than a simple linear relationship. The embodiment of the present application proposes determining the target indicator of the target area through the target indicators corresponding to the multiple blocks respectively, which may include: S21, obtaining the geometric mean of the target indicators corresponding to the multiple blocks respectively; S22, setting the geometric mean as the target indicator of the target area.
[0056] For example, assuming that the block is 16*16 pixels, the target index of the target area is λ SB , can be determined by the following formula 5, where N is the number of blocks, N is greater than or equal to 1, 16 x 16 i is the position of the current block in SB, x, i are greater than or equal to 1:
[0057]
[0058] Before adjusting the initial index of any block among the multiple blocks by the target adjustment factor corresponding to the block to obtain multiple target indexes, the embodiment of the present application further proposes an optional implementation method for initially adjusting the QP, which includes:
[0059] S31, for any block among the multiple blocks, obtaining a first incremental quantization parameter dQP of the block, to obtain multiple first dQPs, wherein the first dQP is determined at least by the complexity of the content of the block and the degree to which the block is referenced by other video frames in the to-be-encoded video frame;
[0060] S32, obtaining a common part of a plurality of first dQPs;
[0061] S33, determining a second frame level QP by using the common part and the first frame level QP corresponding to the video frame to be encoded;
[0062] S34: Determine a plurality of second dQPs by using the common portion and the plurality of first dQPs, wherein the number of bits occupied by the plurality of second dQPs in the video coding stream is less than that of the plurality of first dQPs.
[0063] The first dQP is the initial dQP of the block, and the second dQP is the target dQP of the block after adjustment. The common part of the multiple first dQPs is the same or similar QP area in the video encoding process under different QP variation ranges. The first frame level QP is the initial frame level QP, and the second frame level QP is the target frame level QP after adjustment.
[0064] Optionally, the common part of the above-mentioned multiple first dQPs can be an average value of the multiple first dQPs, or a median value of the multiple first dQPs, or a value determined by performing K-means clustering on the multiple first dQPs. Specifically, a suitable calculation method can be selected according to requirements and content characteristics, so as to achieve better compression efficiency and video quality in the video encoding process.
[0065] Taking the common part as the median of multiple first dQPs as an example, assuming there are 3 blocks, the first dQPs of these 3 blocks are 2, 3, and 5 respectively, the median is 3, the first frame level QP is 25, then the second frame level QP is 28, and the second dQPs of the 3 blocks are -1, 0, and 2 respectively. In video encoding, -1, 0, and 2 occupy less data in the bitstream than 2, 3, and 5.
[0066] Through the above steps S31 to S34, the QP at the frame level is adjusted first, and then the dQP at the block level is determined, so that the dQP at the block level becomes smaller, thereby reducing the amount of data that needs to be transmitted in the bitstream.
[0067] When the common part is an average value of the plurality of first dQPs, determining the second frame-level QP by using the common part and the first frame-level QP corresponding to the to-be-encoded video frame includes: S331, determining a sum value between the average value and the first frame-level QP; S332, determining the sum value as the second frame-level QP. That is, the second frame-level QP is the first frame-level QP plus an average value of the plurality of first dQPs.
[0068] Determining a plurality of second dQPs by using the common part and the plurality of first dQPs includes: S341, determining the differences between the plurality of first dQPs and the average value, respectively, to obtain a plurality of differences; S342, determining the plurality of second dQPs based on the plurality of differences. That is, the second dQP is the first dQP minus the average value of the plurality of first dQPds.
[0069] Exemplarily, assuming that the second frame level QP is QP frame 2. The first frame level QP is QP frame 1, the above average value is Then the above QP frame 2 can be determined by the following formula 6:
[0070]
[0071] The second dQP is determined by the following formula 7 and formula 8, where dQP TPL is the dQP determined by the degree to which the current block is referenced by other video frames in the video frame to be encoded, dQP aq is the dQP determined by the complexity of the current block content:
[0072]
[0073] QDQ sb 1=dQP TPL +dQP aq (Formula 8)
[0074] For example, suppose there are 3 blocks, the first dQP of these 3 blocks are 3, 4, and 5 respectively, the average value is 4, and the first frame level QP is 25, then the second frame level QP can be 29, and the second dQP of these 3 blocks can be -1, 0, and 1 respectively. In video encoding, -1, 0, and 1 occupy less data in the bitstream than 3, 4, and 5.
[0075] After determining multiple second dQPs, the embodiment of the present application further proposes: S41, adjusting the initial variance of any block among the multiple blocks to obtain a target variance; S42, adjusting the second dQP by the target variance. That is, in this optional implementation, the initial variance can be adjusted based on the SSIM indicator, and the second dQP can be adjusted based on the adjusted variance.
[0076] Optionally, adjusting the initial variance of any one of the multiple blocks to obtain a target variance may include: S411, determining a fitting strategy based on the complexity of the block content; S412, determining a fitting relationship between the target variance and the initial variance according to the fitting strategy; S413, adding a second preset offset based on the fitting relationship to determine the target variance, wherein the second preset offset is determined based on the SSIM index and the pixel size of the block.
[0077] The above fitting strategies include but are not limited to: linear fitting, exponential fitting, and segmented fitting. You can choose a suitable fitting strategy according to specific scenarios and requirements to achieve a more refined adjustment effect.
[0078] Taking exponential fitting as an example, the embodiment of the present application proposes the following formula 9 to determine the target variance, where var1 is the initial variance, var2 is the target variance, and C2 is a constant determined by the SSIM indicator:
[0079]
[0080] Optionally, adjusting the second dQP by the target variance may include: S421, determining a smoothing strategy between the second dQP and the target variance; S422, adjusting the second dQP according to the smoothing strategy, wherein the smoothing strategy is used to compress the target variance.
[0081] The above-mentioned smoothing strategies include but are not limited to: tenth root transformation, logarithmic transformation, square root transformation, etc. Each smoothing method has a different degree of compression on the original variance. The appropriate smoothing method can be selected according to the specific application scenario and requirements.
[0082] Taking the tenth root transformation as an example, the embodiment of the present application proposes the following formula 10 to determine the adjusted dQP, that is, qp adjust , where the correction bit width bitdepth correction See Formula 11 to ensure the accuracy and stability of the video encoding algorithm. Wherein, 1.f is a floating point number.
[0083]
[0084] bitdepth correction =1.f / (1<<(2*(bitdepth-8)) (Formula 11)
[0085] In summary, the embodiments of the present application can find out the texture characteristics of different areas based on the variance of the block. For complex areas that are not sensitive to the human eye, the QP can be appropriately increased, and for flat areas, the QP can be reduced. When there are both flat and uneven texture areas inside the SB, the λ of different areas can be further adjusted. For frames at high time domain levels, the bit overhead of dQP can also be saved by adjusting λ. In addition, compared with the optimization of encoder pre-processing in related technologies, the embodiments of the present application can significantly improve encoder performance without increasing computing power overhead.
[0086] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0087] The technical solution of the present application and how the technical solution of the present application solves the above-mentioned technical problems are described in detail below with specific embodiments. The several specific embodiments listed can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present application will be described in detail below with reference to the accompanying drawings.
[0088] Corresponding to the application scenario and method of the method provided in the embodiment of the present application, the embodiment of the present application also provides a video encoding device. Figure 3 The structure block diagram of a video encoding device according to an embodiment of the present application is shown, and the device may include:
[0089] A division module 32 is used for dividing a target area in a super block SB in a video frame to be encoded into a plurality of blocks of a specified pixel size;
[0090] A first adjustment module 34 is used to adjust the initial index of any block among the multiple blocks by using the target adjustment factor corresponding to the block to obtain the target indexes respectively corresponding to the multiple blocks, wherein the target adjustment factor includes at least a first adjustment factor and a second adjustment factor, the first adjustment factor is used to characterize the complexity of the content of the block, and the second adjustment factor is used to characterize the degree to which the block is referenced by other video frames in the video frame to be encoded, and there is a preset relationship between the value of the initial index and the value of the block-level quantization parameter QP;
[0091] The processing module 36 is used to determine the target indicator of the target area according to the target indicators respectively corresponding to the multiple blocks, and encode the target area based on the target indicator of the target area.
[0092] pass Figure 3 The device shown divides a target area in a super block SB in a video frame to be encoded into multiple blocks of specified pixel sizes; for any block in the multiple blocks, an initial index of the block is adjusted by a target adjustment factor corresponding to the block to obtain target indexes respectively corresponding to the multiple blocks, wherein the target adjustment factor includes at least a first adjustment factor and a second adjustment factor, the first adjustment factor is used to characterize the complexity of the content of the block, and the second adjustment factor is used to characterize the degree to which the block is referenced by other video frames in the video frame to be encoded, and there is a preset relationship between the value of the initial index and the value of the block-level quantization parameter QP; the target index of the target area is determined by the target indexes respectively corresponding to the multiple blocks, and the target area is encoded based on the target index of the target area. That is to say, the embodiment of the present application achieves the purpose of fine-tuning QP by jointly adjusting the initial indicators of different areas in SB through a first adjustment factor used to characterize the complexity of block content and a second adjustment factor used to characterize the degree to which the block is referenced by other video frames in the video frame to be encoded, thereby solving the technical problem in the related art that the SB has both flat texture areas and uneven texture areas, and the use of the same QP leads to low encoding quality and encoding efficiency, thereby achieving the technical effect of improving encoding quality and encoding efficiency.
[0093] In one possible implementation, the first adjustment module 34 includes: a first determination unit, used to determine the first adjustment factor corresponding to the block through the variance corresponding to the block; an integration unit, used to integrate the first adjustment factor corresponding to the block and the second adjustment factor corresponding to the block to obtain a target adjustment factor corresponding to the block; and a second determination unit, used to adjust the initial indicator of the block through the target adjustment factor corresponding to the block to obtain the target indicator corresponding to the block.
[0094] The first determination unit includes: an establishment subunit, which is used to establish a mapping relationship between the first adjustment factor and the variance; a determination subunit, which is used to add a first preset offset based on the mapping relationship to determine the first adjustment factor, wherein the first preset offset is determined based on a structural similarity index SSIM indicator. The integration unit is also used to determine the product of the first adjustment factor and the second adjustment factor as the target adjustment factor.
[0095] The processing module 36 includes: a third determining unit, configured to obtain a geometric mean value of target indicators respectively corresponding to the plurality of blocks; and a setting unit, configured to set the geometric mean value as the target indicator of the target area.
[0096] The above-mentioned device also includes: a first determination module, which is used to adjust the initial index of any block among the multiple blocks by the target adjustment factor corresponding to the block to obtain a plurality of target indicators, and then, for any block among the multiple blocks, obtain a first incremental quantization parameter dQP of the block to obtain a plurality of first dQPs, wherein the first dQP is determined at least by the complexity of the content of the block and the degree to which the block is referenced by other video frames in the video frame to be encoded; a second determination module, which is used to obtain a common part of the multiple first dQPs; a third determination module, which is used to determine a second frame-level QP by the common part and the first frame-level QP corresponding to the video frame to be encoded; and a fourth determination module, which is used to determine a plurality of second dQPs by the common part and the plurality of first dQPs, wherein the number of bits occupied by the plurality of second dQPs in the video encoding stream is less than that of the plurality of first dQPs.
[0097] Optionally, the third determination module includes: a first processing unit, configured to determine, when the common portion is an average value of the plurality of first dQPs, a sum value between the average value and the first frame-level QP; and determine the sum value as the second frame-level QP;
[0098] The fourth determination module includes: a second processing unit, configured to determine the differences between the plurality of first dQPs and the average value respectively to obtain a plurality of differences; and determine the plurality of second dQPs based on the plurality of differences.
[0099] The above device also includes: a second adjustment module, used to adjust the initial variance of any block among the multiple blocks after determining the multiple second dQPs to obtain a target variance; and a third adjustment module, used to adjust the second dQP according to the target variance.
[0100] The second adjustment module includes: a fourth determination unit, configured to determine a fitting strategy based on the complexity of the block content; a fifth determination unit, configured to determine a fitting relationship between the target variance and the initial variance according to the fitting strategy; a sixth determination unit, configured to add a second preset offset based on the fitting relationship to determine the target variance, wherein the second preset offset is determined based on the SSIM index and the pixel size of the block;
[0101] The third adjustment module includes: a seventh determination unit, configured to determine a smoothing strategy between the second dQP and the target variance; and an adjustment unit, configured to adjust the second dQP according to the smoothing strategy, wherein the smoothing strategy is used to compress the target variance.
[0102] In summary, the embodiments of the present application can find out the texture characteristics of different areas based on the variance of the block. For complex areas that are not sensitive to the human eye, the QP can be appropriately increased, and for flat areas, the QP can be reduced. When there are both flat and uneven texture areas inside the SB, the λ of different areas can be further adjusted. For frames at high time domain levels, the bit overhead of dQP can also be saved by adjusting λ. In addition, compared with the optimization of encoder pre-processing in related technologies, the embodiments of the present application can significantly improve encoder performance without increasing computing power overhead.
[0103] The functions of each module in each device in the embodiments of the present application can be found in the corresponding description in the above method, and have corresponding beneficial effects, which will not be repeated here.
[0104] Corresponding to the application scenario and method of the method provided in the embodiment of the present application, the embodiment of the present application also provides a video encoder. Figure 4 The structure block diagram of a video encoder according to an embodiment of the present application is shown, and the video encoder may include:
[0105] The pre-processing module 42 is used for encoding using the above-mentioned video encoding method.
[0106] pass Figure 4The video encoder shown, for a super block SB in a video frame to be encoded, divides a target area in the SB into multiple blocks of specified pixel sizes; for any block in the multiple blocks, adjusts the initial index of the block by a target adjustment factor corresponding to the block, and obtains target indexes corresponding to the multiple blocks respectively, wherein the target adjustment factor includes at least a first adjustment factor and a second adjustment factor, the first adjustment factor is used to characterize the complexity of the content of the block, and the second adjustment factor is used to characterize the degree to which the block is referenced by other video frames in the video frame to be encoded, and there is a preset relationship between the value of the initial index and the value of the block-level quantization parameter QP; the target index of the target area is determined by the target indexes corresponding to the multiple blocks respectively, and the target area is encoded based on the target index of the target area. That is to say, the embodiment of the present application achieves the purpose of fine-tuning QP by jointly adjusting the initial indicators of different areas in SB through a first adjustment factor used to characterize the complexity of block content and a second adjustment factor used to characterize the degree to which the block is referenced by other video frames in the video frame to be encoded, thereby solving the technical problem in the related art that the SB has both flat texture areas and uneven texture areas, and the use of the same QP leads to low encoding quality and encoding efficiency, thereby achieving the technical effect of improving encoding quality and encoding efficiency.
[0107] Figure 5 is a block diagram of an electronic device used to implement an embodiment of the present application. Figure 5 As shown, the electronic device includes: a memory 501 and a processor 502. The memory 501 stores a computer program that can be run on the processor 502. When the processor 502 executes the computer program, the method in the above embodiment is implemented. The number of the memory 501 and the processor 502 can be one or more.
[0108] The electronic device also includes:
[0109] The communication interface 503 is used to communicate with external devices and perform data exchange transmission.
[0110] If the memory 501, the processor 502 and the communication interface 503 are implemented independently, the memory 501, the processor 502 and the communication interface 503 can be connected to each other through a bus and communicate with each other. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 5Only one thick line is used in the diagram, but this does not mean that there is only one bus or only one type of bus.
[0111] Optionally, in a specific implementation, if the memory 501, the processor 502 and the communication interface 503 are integrated on a chip, the memory 501, the processor 502 and the communication interface 503 can communicate with each other through an internal interface.
[0112] An embodiment of the present application provides a computer-readable storage medium storing a computer program, which implements the method provided in the embodiment of the present application when the program is executed by a processor.
[0113] An embodiment of the present application also provides a chip, which includes a processor for calling and executing instructions stored in the memory from the memory, so that a communication device equipped with the chip executes the method provided by the embodiment of the present application.
[0114] An embodiment of the present application also provides a chip, including: an input interface, an output interface, a processor and a memory, wherein the input interface, the output interface, the processor and the memory are connected via an internal connection path, and the processor is used to execute the code in the memory. When the code is executed, the processor is used to execute the method provided in the embodiment of the application.
[0115] It should be understood that the processor may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc. It is worth noting that the processor may be a processor supporting the Advanced RISC Machines (ARM) architecture.
[0116] Further, optionally, the above-mentioned memory may include a read-only memory and a random access memory. The memory may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memory. Among them, the non-volatile memory may include a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may include a random access memory (RAM), which is used as an external cache. By way of exemplary but not limiting description, many forms of RAM are available. For example, static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM) and direct memory bus random access memory (DR RAM).
[0117] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function according to the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium.
[0118] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine different embodiments or examples described in this specification and the features of different embodiments or examples, unless they are contradictory.
[0119] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include at least one of the features. In the description of this application, the meaning of "plurality" is two or more, unless otherwise clearly and specifically defined.
[0120] Any process or method described in the flow chart or otherwise described herein can be understood as a module, fragment or portion of a code representing one or more executable instructions for implementing the steps of a specific logical function or process. And the scope of the preferred embodiment of the present application includes other implementations, in which the functions may not be performed in the order shown or discussed, including in a substantially simultaneous manner or in a reverse order according to the functions involved.
[0121] The logic and / or steps described in the flowchart or otherwise described herein, for example, can be considered as an ordered list of executable instructions for implementing logical functions, which can be specifically implemented in any computer-readable medium for use by an instruction execution system, device or apparatus (such as a computer-based system, a system including a processor, or other system that can fetch instructions from an instruction execution system, device or apparatus and execute instructions), or used in combination with these instruction execution systems, devices or apparatuses.
[0122] It should be understood that the various parts of the present application can be implemented with hardware, software, firmware or a combination thereof. In the above embodiments, multiple steps or methods can be implemented with software or firmware stored in a memory and executed by a suitable instruction execution system. All or part of the steps of the above embodiment method can be completed by instructing the relevant hardware through a program, which can be stored in a computer-readable storage medium, and when the program is executed, it includes one of the steps of the method embodiment or a combination thereof.
[0123] In addition, each functional unit in each embodiment of the present application can be integrated into a processing module, or each unit can exist physically separately, or two or more units can be integrated into one module. The above-mentioned integrated module can be implemented in the form of hardware or in the form of a software functional module. If the above-mentioned integrated module is implemented in the form of a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. The storage medium can be a read-only memory, a disk or an optical disk, etc.
[0124] The above is only an exemplary embodiment of the present application, but the protection scope of the present application is not limited thereto. Any technician familiar with the technical field can easily think of various changes or substitutions within the technical scope recorded in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application shall be based on the protection scope of the claims.
Claims
1. A video encoding method, comprising: For a super block SB in a video frame to be encoded, dividing a target area in the SB into a plurality of blocks of a specified pixel size; For any block among the multiple blocks, adjusting the initial index of the block by the target adjustment factor corresponding to the block to obtain the target indexes respectively corresponding to the multiple blocks, wherein the target adjustment factor includes at least a first adjustment factor and a second adjustment factor, the first adjustment factor is used to characterize the complexity of the content of the block, and the second adjustment factor is used to characterize the degree to which the block is referenced by other video frames in the video frame to be encoded, and there is a preset relationship between the value of the initial index and the value of the block-level quantization parameter QP; The target index of the target area is determined by the target indexes respectively corresponding to the multiple blocks, and the target area is encoded based on the target index of the target area.
2. The method according to claim 1, wherein: For any block among the multiple blocks, adjusting the initial index of the block by using the target adjustment factor corresponding to the block to obtain the target index of the block includes: Determine a first adjustment factor corresponding to the block according to the variance corresponding to the block; Integrate a first adjustment factor corresponding to the block and a second adjustment factor corresponding to the block to obtain a target adjustment factor corresponding to the block; The initial index of the block is adjusted by the target adjustment factor corresponding to the block to obtain the target index corresponding to the block.
3. The method according to claim 2, wherein: Determining a first adjustment factor corresponding to the block by the variance corresponding to the block includes: establishing a mapping relationship between the first adjustment factor and the variance; adding a first preset offset based on the mapping relationship to determine the first adjustment factor, wherein the first preset offset is determined based on a structural similarity index SSIM indicator; Integrating a first adjustment factor corresponding to the block and a second adjustment factor corresponding to the block to obtain a target adjustment factor corresponding to the block includes: determining a product of the first adjustment factor and the second adjustment factor as the target adjustment factor.
4. The method according to claim 1, wherein: Determining the target indicator of the target area by using the target indicators respectively corresponding to the multiple blocks includes: Obtaining geometric mean values of target indicators respectively corresponding to the multiple blocks; The geometric mean is set as the target indicator for the target area.
5. The method according to claim 1, wherein: Before adjusting the initial index of any block among the multiple blocks by using the target adjustment factor corresponding to the block to obtain multiple target indexes, the method further includes: For any block among the multiple blocks, obtain a first incremental quantization parameter dQP of the block to obtain multiple first dQPs, wherein the first dQP is determined at least by the complexity of the content of the block and the degree to which the block is referenced by other video frames in the to-be-encoded video frame; Obtaining a common portion of the plurality of first dQPs; Determine a second frame level QP by using the common part and a first frame level QP corresponding to the video frame to be encoded; A plurality of second dQPs are determined by using the common part and the plurality of first dQPs, wherein the number of bits occupied by the plurality of second dQPs in the video coding stream is less than that of the plurality of first dQPs.
6. The method according to claim 5, wherein: When the common portion is an average value of the plurality of first dQPs, Determining the second frame level QP by using the common part and the first frame level QP corresponding to the to-be-encoded video frame comprises: determining a sum value between the average value and the first frame level QP; and determining the sum value as the second frame level QP; Determining a plurality of second dQPs by using the common part and the plurality of first dQPs includes: determining differences between the plurality of first dQPs and the average value to obtain a plurality of differences; and determining the plurality of second dQPs based on the plurality of differences.
7. The method according to claim 6, wherein: After determining the plurality of second dQPs, the method further includes: Adjusting an initial variance of any block among the plurality of blocks to obtain a target variance; The second dQP is adjusted by the target variance.
8. The method according to claim 7, wherein: The initial variance of any one of the plurality of blocks is adjusted to obtain a target variance, comprising: determining a fitting strategy based on the complexity of the block content; determining a fitting relationship between the target variance and the initial variance according to the fitting strategy; adding a second preset offset based on the fitting relationship to determine the target variance, wherein the second preset offset is determined based on an SSIM index and a pixel size of the block; Adjusting the second dQP by the target variance includes: determining a smoothing strategy between the second dQP and the target variance; and adjusting the second dQP according to the smoothing strategy, wherein the smoothing strategy is used to compress the target variance.
9. A video encoder, comprising at least a pre-processing module, wherein: The pre-processing module performs encoding using the method described in any one of claims 1-8.
10. An electronic device comprising a memory, a processor and a computer program stored in the memory, wherein the processor implements the method according to any one of claims 1 to 8 when executing the computer program.
11. A computer-readable storage medium, wherein a computer program is stored in the computer-readable storage medium, and when the computer program is executed by a processor, the method according to any one of claims 1 to 8 is implemented.
12. A computer program product, comprising computer instructions, wherein when the computer instructions are executed by a processor, the method according to any one of claims 1 to 8 is implemented.
Citation Information
Cited By
Video encoding method, video encoder, electronic equipment, storage medium and computer program product
CN120547334A
Video encoding method, video encoder, system and program product
CN120614461A