Video encoding method, device, storage medium and electronic device

By calculating the average SATD value and maximum bit rate of the video frame to obtain the target quantization parameters, the problem of low video encoding accuracy is solved, and stable video transmission and optimization rate distortion performance under limited bandwidth is achieved.

CN114845106BActive Publication Date: 2025-08-15PEKING UNIV SHENZHEN GRADUATE SCHOOL +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110138899.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-02-01
Publication Date
2025-08-15
Estimated Expiration
2041-02-01

AI Technical Summary

Technical Problem

In the prior art, video encoding accuracy is low, and there is a laying hen paradox problem, which leads to prediction errors in bit rate control and affects video playback quality.

Method used

By calculating the average SATD value of the sum of the residual absolute values of each pixel in the video frame to be encoded, the target quantization parameters are obtained in combination with the maximum bit rate, and encoding them according to the target quantization parameters to indicate the output code rate.

Benefits of technology

The accuracy of video encoding is improved, prediction errors are eliminated, and the stability of video transmission and rate distortion performance optimization is achieved under limited bandwidth.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114845106B_ABST
    Figure CN114845106B_ABST
Patent Text Reader

Abstract

The present invention discloses a video encoding method, device, storage medium, and electronic device in a cloud technology scenario, and specifically relates to cloud computing, big data, and other technologies. The method comprises: determining a video frame to be encoded from a video stream to be played; obtaining an average SATD value of the sum of the absolute values of the residuals of each pixel in the video frame to be encoded, and obtaining a target quantization parameter based on the average SATD value and the target number of bits corresponding to the maximum bit rate of the video frame to be encoded, wherein the average SATD value is used to indicate the content complexity of the video frame to be encoded, and the target quantization parameter is used to indicate the output bit rate of the video frame to be encoded; and encoding the video frame to be encoded according to the output bit rate indicated by the target quantization parameter. The present invention solves the technical problem of low video encoding accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computers, and in particular to a video encoding method, device, storage medium, and electronic equipment. Background Art

[0002] With the rapid development of information technology, online video applications are growing rapidly. Its core value is gradually extending downstream in the industry chain, driving the integration of different industries. In 2015, video streams accounted for 85% of the data exchanged on the internet. Compared to image and audio data, video data has a relatively large storage and transmission scale, and faces more challenges in network distribution. In particular, new applications such as ultra-high-definition user experience are limited by technical bottlenecks. Limited by the real-time, variable, and relatively limited network bandwidth, the playback quality of online video services is often unsatisfactory. Rate control is becoming increasingly important in areas such as video image compression and network multimedia transmission. As an algorithmic module on the encoder side, rate control strictly controls the video bitstream rate output on the channel based on the available network bandwidth, achieving stable video transmission and achieving the optimal balance between visual quality and bandwidth utilization. Using reasonable rate control to achieve video encoding has long been a hot topic and a key research focus.

[0003] Existing techniques often use MAD (Mean Absolute Difference) as a measure of sequence texture complexity to perform relevant calculations for video encoding. However, because MAD is the absolute error between the reconstructed and original pixels, it cannot be obtained before the current block is encoded. Rate control requires quantization parameters calculated based on the MAD value. However, the MAD calculation, after the RDO process, requires quantization parameters, leading to a chicken-and-egg paradox. This, in turn, leads to prediction errors in rate control, thus affecting video encoding accuracy. Consequently, existing techniques suffer from low video encoding accuracy.

[0004] To address the above-mentioned problems, no effective solutions have been proposed so far. Summary of the Invention

[0005] Embodiments of the present invention provide a video encoding method, apparatus, storage medium, and electronic device to at least solve the technical problem of low video encoding accuracy.

[0006] According to one aspect of an embodiment of the present invention, a video encoding method is provided, comprising: determining a video frame to be encoded from a video stream to be played; upon obtaining an average SATD value of the sum of the absolute values of the residuals of each pixel in the above-mentioned video frame to be encoded, obtaining a target quantization parameter based on the above-mentioned average SATD value and a target number of bits corresponding to the maximum bit rate of the above-mentioned video frame to be encoded, wherein the above-mentioned average SATD value is used to indicate the content complexity of the above-mentioned video frame to be encoded, and the above-mentioned target quantization parameter is used to indicate the output bit rate of the above-mentioned video frame to be encoded; and encoding the above-mentioned video frame to be encoded according to the output bit rate indicated by the above-mentioned target quantization parameter.

[0007] According to another aspect of an embodiment of the present invention, a video encoding device is also provided, including: a determination unit, used to determine a video frame to be encoded from a video stream to be played; an acquisition unit, used to obtain a target quantization parameter based on the average SATD value of the sum of the absolute values of the residuals of each pixel in the above-mentioned video frame to be encoded and the target number of bits corresponding to the maximum bit rate of the above-mentioned video frame to be encoded, wherein the above-mentioned average SATD value is used to indicate the content complexity of the above-mentioned video frame to be encoded, and the above-mentioned target quantization parameter is used to indicate the output bit rate of the above-mentioned video frame to be encoded; and an encoding unit, used to encode the above-mentioned video frame to be encoded according to the output bit rate indicated by the above-mentioned target quantization parameter.

[0008] According to another aspect of the embodiments of the present invention, a computer-readable storage medium is provided, in which a computer program is stored. The computer program is configured to execute the above-mentioned video encoding method when running.

[0009] According to another aspect of an embodiment of the present invention, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the video encoding method through the computer program.

[0010] In an embodiment of the present invention, a video frame to be encoded is determined from a video stream to be played; when the average SATD value of the sum of the absolute values of the residuals of each pixel in the above-mentioned video frame to be encoded is obtained, a target quantization parameter is obtained according to the above-mentioned average SATD value and the target number of bits corresponding to the maximum bit rate of the above-mentioned video frame to be encoded, wherein the above-mentioned average SATD value is used to indicate the content complexity of the above-mentioned video frame to be encoded, and the above-mentioned target quantization parameter is used to indicate the output bit rate of the above-mentioned video frame to be encoded; the above-mentioned video frame to be encoded is encoded according to the output bit rate indicated by the above-mentioned target quantization parameter, and by using the average SATD value that can be obtained without encoding as a measure of the content complexity of the video frame to be encoded, the chicken and egg paradox problem existing in the prior art is overcome, the prediction error is eliminated, and the technical purpose of improving the accuracy of the output bit rate indicated by the target quantization parameter is achieved, thereby achieving the technical effect of improving the encoding accuracy of the video, and solving the technical problem of low encoding accuracy of the video. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of this application. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:

[0012] Figure 1 is a schematic diagram of an application environment of an optional video encoding method according to an embodiment of the present invention;

[0013] Figure 2 is a schematic diagram of a flowchart of an optional video encoding method according to an embodiment of the present invention;

[0014] Figure 3 is a schematic diagram of an optional video encoding method according to an embodiment of the present invention;

[0015] Figure 4 is a schematic diagram of another optional video encoding method according to an embodiment of the present invention;

[0016] Figure 5 is a schematic diagram of another optional video encoding method according to an embodiment of the present invention;

[0017] Figure 6 is a schematic diagram of an optional video encoding device according to an embodiment of the present invention;

[0018] Figure 7 is a schematic diagram of another optional video encoding device according to an embodiment of the present invention;

[0019] Figure 8 is a schematic diagram of another optional video encoding device according to an embodiment of the present invention;

[0020] Figure 9 FIG. 4 is a schematic structural diagram of an optional electronic device according to an embodiment of the present invention. DETAILED DESCRIPTION

[0021] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.

[0022] It should be noted that the terms "first", "second", etc. in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0023] Cloud technology refers to a hosting technology that unifies hardware, software, network and other resources within a wide area network or local area network to achieve data computing, storage, processing and sharing.

[0024] Cloud technology is a general term for network technologies, information technologies, integration technologies, management platform technologies, and application technologies based on the cloud computing business model. It can form a resource pool that can be used flexibly and conveniently on demand. Cloud computing technology will become a crucial support. Backend services for technical network systems, such as video websites, image websites, and more portals, require extensive computing and storage resources. With the rapid development and application of the internet industry, every item will likely have its own unique identifier, requiring transmission to backend systems for logical processing. Different levels of data will be processed separately, and data from various industries will require a strong system backend, which can only be achieved through cloud computing.

[0025] Cloud computing is a computing model that distributes computing tasks across a resource pool consisting of a large number of computers, enabling various application systems to access computing power, storage space, and information services as needed. The network that provides these resources is called the "cloud." To users, these resources appear infinitely scalable and can be accessed at any time, used on demand, expanded at any time, and paid for on a per-use basis.

[0026] As a provider of cloud computing infrastructure, a cloud computing resource pool (referred to as a cloud platform, generally referred to as an IaaS (Infrastructure as a Service) platform) is established. Various types of virtual resources are deployed in the resource pool for external customers to choose and use. The cloud computing resource pool mainly includes: computing devices (virtualized machines, including operating systems), storage devices, and network devices.

[0027] Based on logical functional divisions, the PaaS (Platform as a Service) layer can be deployed on top of the IaaS (Infrastructure as a Service) layer, and the SaaS (Software as a Service) layer can be deployed on top of the PaaS layer. SaaS can also be deployed directly on top of IaaS. PaaS is a platform for software execution, such as databases and web containers. SaaS is a variety of business software, such as web portals and text messaging apps. Generally speaking, SaaS and PaaS are upper layers relative to IaaS.

[0028] According to one aspect of an embodiment of the present invention, a video encoding method is provided. Optionally, as an optional implementation, the video encoding method can be applied to, but is not limited to, Figure 1 In the environment shown, it may include, but is not limited to, a user device 102, a network 110, and a server 112. The user device 102 may include, but is not limited to, a display 108, a processor 106, and a memory 104.

[0029] The specific process can be as follows:

[0030] In step S102, the user device 102 determines a video frame to be encoded from a video stream to be played. The video stream to be played may be a video stream to be played on a display of another user device or a video stream to be played on the display 108 of the user device 102, which is not limited here.

[0031] In steps S104-S106, the user device 102 sends the video frame to be encoded to the server 112 via the network 110, wherein the video frame to be encoded may be, but is not limited to, a group of video frames or a single video frame, which is not limited here;

[0032] In step S108, the server 112 calculates, through the processing engine 116, an average SATD value of the sum of the absolute values of the residuals of each pixel in the to-be-encoded video frame and a target number of bits corresponding to the maximum bit rate, further calculates a target quantization parameter based on the average SATD value and the target number of bits, and encodes the to-be-encoded video frame using the output bit rate indicated by the target quantization parameter, thereby generating a target video frame and completing encoding of the to-be-encoded video frame.

[0033] In steps S110-S112, assuming that the video stream to be played is a video stream to be played on the display 108 of another user device, the server 112 sends the target video frame to the user device 102 via the network 110. The processor 106 in the user device 102 displays the target video frame on the display 108 and stores the target video frame in the memory 104. The adjustment and determination results may be stored in, but are not limited to, the server 112 or the user device 102.

[0034] remove Figure 1 In addition to the examples shown, the above steps can be independently completed by the user device 102. That is, the user device 102 performs the steps of calculating the average SATD value of the sum of the residual absolute values of each pixel in the video frame to be encoded, calculating the target number of bits corresponding to the maximum bit rate, obtaining the target quantization parameter, and encoding the video frame to be encoded, thereby reducing the processing pressure on the server. The user device 102 includes but is not limited to a handheld device (such as a mobile phone), a laptop computer, a desktop computer, an in-vehicle device, etc. The present invention does not limit the specific implementation of the user device 102.

[0035] In addition, the server can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The terminal can be a smartphone, tablet computer, laptop computer, desktop computer, smart speaker, smart watch, etc., but is not limited to these. The terminal and the server can be directly or indirectly connected via wired or wireless communication, and this application does not impose any restrictions on this.

[0036] Alternatively, as an optional implementation, Figure 2 As shown, the video encoding method includes:

[0037] S202, determining a video frame to be encoded from the video stream to be played;

[0038] S204: When an average SATD value of the sum of the residual absolute values of each pixel in the video frame to be encoded is obtained, a target quantization parameter is obtained based on the average SATD value and a target number of bits corresponding to the maximum bit rate of the video frame to be encoded, wherein the average SATD value is used to indicate the content complexity of the video frame to be encoded, and the target quantization parameter is used to indicate the output bit rate of the video frame to be encoded;

[0039] S206 , encoding the video frame to be encoded according to the output bit rate indicated by the target quantization parameter.

[0040] It should be noted that the above Figure 2 The video encoding method shown can be used for, but is not limited to, Figure 1 In the video encoder shown in FIG. , the video encoder interacts with other components to complete the encoding process of the video frame to be encoded.

[0041] Optionally, in this embodiment, the above-mentioned video encoding method can be applied to, but not limited to, full I-frame rate control for the AVS3 video coding standard. It can also be applied to, but not limited to, video encoders, including software encoders and hardware encoders, to achieve stability and optimal rate-distortion performance of video transmission tasks under limited bandwidth constraints. Specifically, the quantization parameter (QP) of the I-frame is determined based on the SATDpp and target bits of the video frame, and on this basis, the output bit rate indicated by the quantization parameter of the I-frame is proposed to encode the video frame to be encoded, wherein the quantization parameter can be used to determine, but not limited to, the quantization step size.

[0042] Optionally, in this embodiment, the above-mentioned video encoding method can also be applied to application scenarios such as video playback applications, video sharing applications or video conversation applications, but is not limited to it. The videos transmitted in the above-mentioned application scenarios may include but are not limited to: long videos and short videos. For example, long videos can be played episodes with a long playback time (for example, a playback time greater than 10 minutes), or pictures displayed in long video conversations. Short videos can be voice messages for interaction between two or more parties, or videos with a short playback time (for example, a playback time less than or equal to 30 seconds) for display on a sharing platform. The above is only an example. The video encoding method provided in this embodiment can be applied to, but is not limited to, playback devices for playing videos in the above-mentioned application scenarios. After obtaining the video frames to be encoded from the video stream to be played, the target quantization parameters of each video frame to be encoded are determined based on the average SATD value with low complexity to avoid the generation of prediction errors and ensure the accuracy of video encoding. Optionally, it is also possible to, but not limited to, perform bit allocation and QP determination on the three YUV channels separately.

[0043] Optionally, in this embodiment, in the video transmission task, the limited channel bandwidth is difficult to meet the growing demand for video resolution and variety. In order to stably and efficiently transmit the video bit stream, it is necessary to perform rate control at the encoding end. The rate control can be, but is not limited to, by selecting a series of quantization parameters (including frame level and block level), so as to achieve the stability of the video transmission task and the optimal rate-distortion performance under the condition of limited bandwidth constraints. Then, the goal of the rate control work is equivalent to obtaining the quantization parameters by bit allocation and estimation model parameters under the constraint of limited bandwidth, so as to minimize the overall distortion. Specifically, the video frame to be encoded is encoded according to the output rate indicated by the obtained quantization parameter, which can be understood by referring to, but not limited to, the mathematical expression of the following formula (1):

[0044]

[0045] Optionally, in this embodiment, the video frame to be encoded can be a video frame captured in real time, or a video frame corresponding to a stored video. The video frame to be encoded can be an input video frame in an input video frame sequence; the video frame to be encoded can also be a video frame obtained by processing an input video frame in an input video frame sequence according to a corresponding processing method, wherein, after the encoding end processes the input video frame according to the corresponding processing method, the resolution of the video frame to be encoded is smaller than the resolution of the original input video frame. For example, the input video frame can be downsampled according to a corresponding sampling ratio to obtain the video frame to be encoded.

[0046] Specifically, the encoding end can determine the processing method of the input video frame, and process the input video frame according to the processing method of the input video frame to obtain the video frame to be encoded. The processing methods include downsampling processing method and full-resolution processing method. The downsampling processing method refers to downsampling the input video frame to obtain the video frame to be encoded, and then encoding the obtained video frame to be encoded. The downsampling method in the downsampling processing method can be customized according to needs, including vertical downsampling, horizontal downsampling, vertical and horizontal downsampling, and downsampling can be performed using direct averaging, filters, bicubic interpolation, bilinear interpolation and other algorithms. The full-resolution processing method refers to directly using the input video frame as the video frame to be encoded, and directly encoding the video frame to be encoded based on the original resolution of the input video frame.

[0047] Optionally, in this embodiment, the encoding end can obtain the processing method corresponding to the input video frame based on the current encoding information corresponding to the input video frame and at least one of the image feature information. The current encoding information refers to the video compression parameter information obtained when the video is encoded, such as one or more of the frame type, motion vector, quantization parameter, video source, bit rate, frame rate and resolution. Image feature information refers to information related to the image content, including one or more of the image motion information and image texture information, such as edges. The current encoding information and image feature information reflect the scene, detail complexity or motion intensity corresponding to the video frame. For example, the motion scene can be judged by one or more of the motion vector, quantization parameter or bit rate. A large quantization parameter generally indicates an intense motion, and a large motion vector indicates that the image scene is a large motion scene.

[0048] The average SATD value of the sum of the absolute residual values of each pixel in the video frame to be encoded is used as the representation of the image feature information. Since the average SATD value has low complexity, it can be obtained without encoding. Then, the average SATD value of the sum of the absolute residual values of each pixel in the video frame to be encoded is used to determine the texture complexity of the video frame to be encoded, eliminating the prediction error caused by the chicken-and-egg paradox problem and improving the accuracy of subsequent encoding-related calculations.

[0049] Optionally, in this embodiment, the target number of bits may be, but is not limited to, an expected number of bits obtained by calculating based on the maximum bit rate of the video frame to be encoded, wherein the bit rate (Bitrate) may be, but is not limited to, also known as the bit rate, represented by the letter r or R, indicating the rate of the bit stream after video encoding, and commonly used units are bps, Kbps, Mbps, and Gbps;

[0050] For example, in order to avoid bandwidth congestion or waste, the sliding window idea is used to determine the target bit value of the video frame to be encoded. Specifically, refer to the following formula (2): the target bit number of the video frame to be encoded is determined by the number of remaining bits bits left , remaining frames left , the average target number of bits per frame target And the sliding window (Slide Window, referred to as SW) is determined, where it is assumed that the sliding window is set to 4.

[0051]

[0052] Optionally, in this embodiment, the processing method corresponding to the input video frame can be set according to actual needs. For example, processing parameters corresponding to the input video frame can be obtained, and the corresponding processing method can be obtained based on the processing parameters. The processing parameters are parameters used to determine the processing method, and the specific processing method used can be set according to needs. For example, the processing parameters may include current encoding information and / or image features corresponding to the input video frame.

[0053] In this embodiment, the sum of the absolute residual values of each pixel in the video frame to be encoded can be, but is not limited to, the sum of the absolute values of the residual absolute values after the Hadamard transformation (Sum of Absolute Transformed Difference, SATD), which is used to represent the complexity of the video frame and the coding block content. The average SATD value can be, but is not limited to, the average SATD value per pixel, i.e., SATD per pixel.

[0054] It should be noted that, a video frame to be encoded is determined from a video stream to be played; when the average SATD value of the sum of the absolute values of the residuals of each pixel in the video frame to be encoded is obtained, a target quantization parameter is obtained based on the average SATD value and the target number of bits corresponding to the maximum bit rate of the video frame to be encoded, wherein the average SATD value is used to indicate the content complexity of the video frame to be encoded, and the target quantization parameter is used to indicate the output bit rate of the video frame to be encoded; the video frame to be encoded is encoded according to the output bit rate indicated by the target quantization parameter.

[0055] Further examples illustrate that the optional application scenarios of the above video encoding method are as follows: Figure 3As shown, the video stream 302 to be played is input into the encoding end 304, and the encoding end 304 determines the video frame to be encoded in the video stream 302, and obtains the target quantization parameter according to the average SATD value and the target number of bits corresponding to the maximum bit rate of the video frame to be encoded, and encodes the video frame to be encoded according to the output bit rate indicated by the target quantization parameter to generate an encoded target video frame; the target video frame is packaged and sent to the playback end 306, and the playback end 306 processes the target video frame and plays the processed target video frame on the display screen of the playback end 306.

[0056] Through the embodiments provided by the present application, a video frame to be encoded is determined from a video stream to be played; when the average SATD value of the sum of the absolute values of the residuals of each pixel in the video frame to be encoded is obtained, a target quantization parameter is obtained based on the average SATD value and the target number of bits corresponding to the maximum bit rate of the video frame to be encoded, wherein the average SATD value is used to indicate the content complexity of the video frame to be encoded, and the target quantization parameter is used to indicate the output bit rate of the video frame to be encoded; the video frame to be encoded is encoded according to the output bit rate indicated by the target quantization parameter, and by using the average SATD value that can be obtained without encoding as a measure of the content complexity of the video frame to be encoded, the chicken-and-egg paradox problem existing in the prior art is overcome, the prediction error is eliminated, and the technical purpose of improving the accuracy of the output bit rate indicated by the target quantization parameter is achieved, thereby achieving the technical effect of improving the encoding accuracy of the video.

[0057] As an optional solution, a target quantization parameter is obtained according to the average SATD value and the target number of bits corresponding to the maximum bit rate of the video frame to be encoded, including:

[0058] S1, calculating a target number of bits and an average SATD value to obtain a frame-level quantization parameter, wherein the frame-level quantization parameter is used to indicate a frame-level output bit rate of a video frame to be encoded;

[0059] S2, determining the frame-level quantization parameter as the target quantization parameter; or,

[0060] S3, when a frame-level quantization parameter is obtained by calculating the target number of bits and the average SATD value, obtaining a block-level quantization parameter based on the frame-level quantization parameter, wherein the block-level quantization parameter is used to indicate a block-level output bit rate of the video frame to be encoded;

[0061] S4: Determine the block-level output bit rate as a target quantization parameter.

[0062] Optionally, in this embodiment, the target quantization parameter may be, but is not limited to, indicating different levels of rate control, such as frame-level rate control, block-level rate control, etc. Specifically, the object of frame-level rate control is a coding unit (CodingUnit, CU for short), and the object of block-level rate control is a largest coding unit (LCU for short). Furthermore, the quantization parameter may be, but is not limited to, also divided into different levels of quantization parameters, such as a frame-level quantization parameter for frame-level rate control, a block-level quantization parameter for block-level rate control, etc.

[0063] It should be noted that block-level code control is proposed on the basis of frame-level code control. For example, when the frame-level quantization parameter is obtained by calculating the target number of bits and the average SATD value, the block-level quantization parameter is obtained based on the frame-level quantization parameter; the block-level output bit rate is determined as the target quantization parameter. Compared with the frame-level quantization parameter, the block-level quantization parameter can adaptively compensate for the error of frame-level code control, and can also realize quantization parameter adjustment for subjective quality optimization, thereby further improving the accuracy of rate control and coding performance.

[0064] Through the embodiments provided in the present application, a target number of bits and an average SATD value are calculated to obtain a frame-level quantization parameter, wherein the frame-level quantization parameter is used to indicate the frame-level output bit rate of the video frame to be encoded; the frame-level quantization parameter is determined as the target quantization parameter; or, when the frame-level quantization parameter is obtained by calculating the target number of bits and the average SATD value, a block-level quantization parameter is obtained based on the frame-level quantization parameter, wherein the block-level quantization parameter is used to indicate the block-level output bit rate of the video frame to be encoded; the block-level output bit rate is determined as the target quantization parameter, thereby achieving the purpose of adaptively compensating for frame-level code control errors and adjusting the quantization parameters for subjective quality optimization, thereby realizing the effect of improving the accuracy of bit rate control and the coding performance.

[0065] As an optional solution, the target number of bits and the average SATD value are calculated to obtain the frame-level quantization parameter, including:

[0066] S1, when the video frame to be encoded is a first video frame, calculating a target number of bits and an average SATD value using a target mathematical formula to obtain a frame-level quantization parameter, wherein the target mathematical formula is used to represent a mathematical relationship between the average SATD value, the target number of bits, and the frame-level quantization parameter; or,

[0067] S2, when the video frame to be encoded is a second video frame, updating calculation parameters in the target mathematical formula based on the number of bits consumed by the first video frame and the target number of bits, wherein the second video frame is a next video frame of the first video frame;

[0068] S3, using the updated target mathematical formula to calculate the target number of bits and the average SATD value to obtain a frame-level quantization parameter.

[0069] It should be noted that, when the video frame to be encoded is the first video frame, the target number of bits and the average SATD value are calculated using the target mathematical formula to obtain the frame-level quantization parameter, wherein the target mathematical formula is used to represent the mathematical relationship between the average SATD value, the target number of bits, and the frame-level quantization parameter; or, when the video frame to be encoded is the second video frame, the calculation parameters in the target mathematical formula are updated based on the number of bits consumed by the first video frame and the target number of bits, wherein the second video frame is the next video frame of the first video frame; the target number of bits and the average SATD value are calculated using the updated target mathematical formula to obtain the frame-level quantization parameter. Optionally, the target mathematical formula can be, but is not limited to, a frame-level R(SATDpp, QP) model, an R(QP, G) model, or a URQ model. Mathematical relationships can be, but are not limited to, used to represent the connection and properties between the elements in the formula, such as binary relationships, proportional relationships, linear relationships, etc.

[0070] As a further example, the target mathematical formula may be optionally described as a frame-level R (SATDpp, QP) model, wherein the frame-level R (SATDpp, QP) model is used to represent the relationship between the number of consumed bits of a video frame to be encoded and the content complexity and quantization parameter QP of the frame, wherein the content complexity of the frame is represented by SATDpp, and the initial model parameters in the frame-level R (SATDpp, QP) model are fixed values, but can be continuously updated in subsequent calculations based on the error between the number of encoded bits of the previous frame and the target number of bits, as shown in the following formula (3):

[0071]

[0072] Among them, bpp is used to represent the ratio characteristics of each pixel in the video frame to be encoded, which is used to reflect the bit rate and can be obtained based on the target number of bits. QP is the quantization parameter, SATDpp is used to characterize the content complexity of the video frame to be encoded, and a, b, c, and d are model parameters.

[0073] Furthermore, based on the above frame-level R(SATDpp, QP) model, the corresponding quadratic equation is obtained by transformation, as shown in the following formula (4):

[0074] A.QP 2 +B·QP+C=0 (4);

[0075] Where A = (bpp-d) / SATDpp-d, B = -b, and C = -c. Therefore, the QP can be calculated based on the target bits of the frame and SATDpp, as shown in the following formula (5):

[0076]

[0077] In addition, A is always less than 0, B is always less than 0, and C is always less than 0, so there is a possibility that the above formula (5) has no solution. In this case, let QP = -B / 2A. Then, when a solution exists for formula (5), select the right value as the final QP.

[0078] Furthermore, for non-first video frames (e.g., I-frames), the frame-level RQ model parameters a, b, c, and d are continuously updated by the error between the number of coded bits of the previous frame and the target number of bits. The update formulas are shown in the following formulas (6), (7), (8), and (9):

[0079] a=a-α a bpp error ·SATDpp (6);

[0080]

[0081]

[0082] d=d-α d bpp error (9);

[0083] Wherein, αa, αb, αc, and αd are learning rates for updating model parameters, which can be assigned values of, but not limited to, 0.0001, 0.1, 10, and 0.001, respectively;

[0084] The bit error bpperror is calculated by the following formula (10):

[0085] bpp error =bpp target -bpp real (10);

[0086] Through the embodiments provided by the present application, when the video frame to be encoded is the first video frame, the target number of bits and the average SATD value are calculated using the target mathematical formula to obtain the frame-level quantization parameter, wherein the target mathematical formula is used to represent the mathematical relationship between the average SATD value, the target number of bits and the frame-level quantization parameter; or, when the video frame to be encoded is the second video frame, the calculation parameters in the target mathematical formula are updated based on the number of bits consumed by the first video frame and the target number of bits, wherein the second video frame is the next video frame of the first video frame; the target number of bits and the average SATD value are calculated using the updated target mathematical formula to obtain the frame-level quantization parameter, thereby achieving the purpose of clearly characterizing the mathematical relationship between the average SATD value, the target number of bits and the frame-level quantization parameter using a specific mathematical formula, and realizing the effect of improving the calculation efficiency of the quantization parameter.

[0087] As an optional solution, obtaining the block-level quantization parameter based on the frame-level quantization parameter includes:

[0088] When a reference quantization parameter of a video frame to be encoded is obtained, the frame-level quantization parameter is updated using the reference quantization parameter to obtain a block-level quantization parameter, wherein the reference quantization parameter is used to adjust a rate error of the frame-level quantization parameter.

[0089] Optionally, in this embodiment, the reference quantization parameter can be used for, but is not limited to, feedback adjustment of the frame-level quantization parameter. For example, due to the inaccuracy of the RQ model and the large change in bit rate corresponding to the QP change, there is a certain deviation in the frame-level rate control algorithm, and at the CTU level, the bit error of the previous frame is used.

[0090] It should be noted that, once the reference quantization parameters of the video frame to be encoded are obtained, the frame-level quantization parameters are updated using the reference quantization parameters to obtain the block-level quantization parameters. This feedback adjustment method is used to reduce bitrate error and improve encoding performance. Furthermore, by setting different QPs for LCUs of varying content complexity, subjective encoding quality can be improved to a certain extent.

[0091] Through the embodiments provided in the present application, when the reference quantization parameter of the video frame to be encoded is obtained, the frame-level quantization parameter is updated using the reference quantization parameter to obtain the block-level quantization parameter, wherein the reference quantization parameter is used to adjust the bit rate error of the frame-level quantization parameter, thereby achieving the effect of reducing the bit rate error and improving the encoding performance.

[0092] As an optional solution, the frame-level quantization parameter is updated using the reference quantization parameter to obtain the block-level quantization parameter, including:

[0093] S1, when the video frame to be encoded is the third video frame, determining the frame-level quantization parameter as the block-level quantization parameter; or,

[0094] S2, when the video frame to be encoded is the fourth video frame, updating the reference quantization parameter based on the number of bits consumed by the third video frame;

[0095] S3, updating the frame-level quantization parameter using the updated reference quantization parameter;

[0096] S4, determining the updated frame-level quantization parameter as the block-level quantization parameter.

[0097] Optionally, in this embodiment, when the video frame to be encoded is the first frame, the QP of all LCUs of the video frame to be encoded is equal to the frame-level QP and does not need to be updated; when the video frame to be encoded is not the first frame, the LCU-level QP of the video frame to be encoded will be adjusted according to the relative bit error of the previous frame.

[0098] It should be noted that, when the video frame to be encoded is the third video frame, the frame-level quantization parameter is determined as the block-level quantization parameter; or, when the video frame to be encoded is the fourth video frame, the reference quantization parameter is updated based on the number of bits consumed by the third video frame; the frame-level quantization parameter is updated using the updated reference quantization parameter; and the updated frame-level quantization parameter is determined as the block-level quantization parameter.

[0099] To further illustrate, the optional Figure 4 The specific steps are as follows:

[0100] S402, obtaining a video frame to be encoded;

[0101] S404, determining whether the video frame to be encoded is the first frame, if not, executing S406, if yes, executing S408;

[0102] S406, determining the frame-level quantization parameter as a block-level quantization parameter;

[0103] S408, calculating the number of bits consumed by the video frame to be encoded;

[0104] S410, updating the reference quantization parameter;

[0105] S412, updating the frame-level quantization parameter using the updated reference quantization parameter;

[0106] S414: Determine the updated frame-level quantization parameter as the block-level quantization parameter.

[0107] Through the embodiments provided in the present application, when the video frame to be encoded is the third video frame, the frame-level quantization parameter is determined as the block-level quantization parameter; or, when the video frame to be encoded is the fourth video frame, the reference quantization parameter is updated based on the number of bits consumed by the third video frame; the frame-level quantization parameter is updated using the updated reference quantization parameter; and the updated frame-level quantization parameter is determined as the block-level quantization parameter, thereby achieving the purpose of flexibly updating the quantization parameter and realizing the effect of improving the flexibility of obtaining the quantization parameter.

[0108] As an optional solution, updating the reference quantization parameter based on the number of bits consumed by the third video frame includes at least one of the following:

[0109] S1, when the number of bits consumed by the third video frame does not reach the target number of bits, updating the reference quantization parameter to a first reference quantization parameter, wherein the first reference quantization parameter is used to indicate a frame-level quantization parameter whose average SATD value is greater than or equal to a complexity threshold to be lowered;

[0110] S2, when the number of bits consumed by the third video frame does not reach the target number of bits, updating the reference quantization parameter to a second reference quantization parameter, wherein the second reference quantization parameter is used to indicate an increase in the frame-level quantization parameter whose average SATD value is less than the complexity threshold.

[0111] Optionally, in this embodiment, if the actual bit consumption of the previous frame is less than the target bit, the QP of the (NQP-n) LCUs with larger texture complexity (which can be expressed by the average SATD value) in the video frame to be encoded is reduced; if the actual bit consumption of the previous frame is greater than the target bit, the QP of the (NQP+n) LCUs with smaller texture complexity in the video frame to be encoded is increased, where n is greater than or equal to 0.

[0112] Among them, NQP-1 and NQP+1 are determined by the following formulas (10), (11), and (12), respectively, where E represents the relative bit error, NLCUs represents the number of LCUs in a frame, and κ is the expansion factor. The number of LCUs for QP adjustment is constrained within an appropriate range. The texture complexity of LCUs is calculated based on the gradient method, as shown in the following formula (12):

[0113]

[0114] N QP-1 =E·κ·N LCUs (11);

[0115] N QP+1 =-E·κ·N LCUs (12);

[0116]

[0117] It should be noted that, when the number of bits consumed by the third video frame does not reach the target number of bits, the updated reference quantization parameter is the first reference quantization parameter, wherein the first reference quantization parameter is used to indicate the frame-level quantization parameter to be lowered so that the average SATD value is greater than or equal to the complexity threshold; when the number of bits consumed by the third video frame does not reach the target number of bits, the updated reference quantization parameter is the second reference quantization parameter, wherein the second reference quantization parameter is used to indicate the frame-level quantization parameter to be raised so that the average SATD value is less than the complexity threshold.

[0118] For further example, the optional video frame to be encoded is an I frame, and the update process of its quantization parameter is as follows: Figure 5 The specific steps are as follows:

[0119] S502, obtaining a video frame to be encoded;

[0120] S504, determining whether the video frame to be encoded is the first I frame, if not, executing S506, if yes, executing S508;

[0121] S506: The QP of all LCUs is always equal to the frame-level QP, i.e., no update or adjustment is performed;

[0122] S508, calculating the gradient wells of LCUs and ranking them;

[0123] S510, calculating a relative bit error;

[0124] S512, determining whether the relative bit error is higher than a preset threshold, if not, executing S514, if yes, executing S516;

[0125] S514, lowering the frame-level quantization parameter whose texture complexity is greater than the complexity threshold;

[0126] S516: Increase the frame-level quantization parameter for which the texture complexity is greater than the complexity threshold.

[0127] Through the embodiment provided by the present application, when the number of bits consumed by the third video frame does not reach the target number of bits, the reference quantization parameter is updated to the first reference quantization parameter, wherein the first reference quantization parameter is used to indicate a frame-level quantization parameter whose average SATD value is greater than or equal to the complexity threshold. When the number of bits consumed by the third video frame does not reach the target number of bits, the reference quantization parameter is updated to the second reference quantization parameter, wherein the second reference quantization parameter is used to indicate a frame-level quantization parameter whose average SATD value is less than the complexity threshold. This achieves the purpose of flexibly updating the quantization parameter and realizes the effect of improving the flexibility of obtaining the quantization parameter.

[0128] As an optional solution, a target quantization parameter is obtained according to the average SATD value and the target number of bits corresponding to the maximum bit rate of the video frame to be encoded, including:

[0129] S1, when a previous frame-level quantization parameter corresponding to a previous video frame of the video frame to be encoded is obtained, determining an upper limit quantization parameter and a lower limit quantization parameter of a target quantization parameter, wherein the upper limit quantization parameter is the sum of the previous frame-level quantization parameter and a preset threshold, and the lower limit quantization parameter is the difference between the previous frame-level quantization parameter and the preset threshold, and the preset threshold is used to constrain the adjustment range of the quantization parameter;

[0130] S2, when the target quantization parameter is greater than or equal to the upper limit quantization parameter, using the upper limit quantization parameter as the target quantization parameter;

[0131] S3: When the target quantization parameter is less than or equal to the lower limit quantization parameter, use the lower limit quantization parameter as the target quantization parameter.

[0132] Optionally, in this embodiment, to reduce bit fluctuation, the difference between the QPs of two adjacent frames may be constrained, but is not limited to constraining the difference between the QPs of two adjacent frames, for example, the difference between the QPs of two adjacent frames is less than or equal to 2.

[0133] It should be noted that, when the previous frame-level quantization parameter corresponding to the previous video frame of the video frame to be encoded is obtained, the upper limit quantization parameter and the lower limit quantization parameter of the target quantization parameter are determined, wherein the upper limit quantization parameter is the sum of the previous frame-level quantization parameter and a preset threshold value, and the lower limit quantization parameter is the difference between the previous frame-level quantization parameter and the preset threshold value, and the preset threshold value is used to constrain the adjustment range of the quantization parameter; when the target quantization parameter is greater than or equal to the upper limit quantization parameter, the upper limit quantization parameter is used as the target quantization parameter; when the target quantization parameter is less than or equal to the lower limit quantization parameter, the lower limit quantization parameter is used as the target quantization parameter.

[0134] Through the embodiments provided by the present application, when the previous frame-level quantization parameter corresponding to the previous video frame of the video frame to be encoded is obtained, the upper limit quantization parameter and the lower limit quantization parameter of the target quantization parameter are determined, wherein the upper limit quantization parameter is the sum of the previous frame-level quantization parameter and a preset threshold, and the lower limit quantization parameter is the difference between the previous frame-level quantization parameter and the preset threshold, and the preset threshold is used to constrain the adjustment range of the quantization parameter; when the target quantization parameter is greater than or equal to the upper limit quantization parameter, the upper limit quantization parameter is used as the target quantization parameter; when the target quantization parameter is less than or equal to the lower limit quantization parameter, the lower limit quantization parameter is used as the target quantization parameter, thereby achieving the effect of reducing bit fluctuation.

[0135] As an optional solution, according to the test conditions established by Audio Video Coding Standard 3.0 (AVS3), it is assumed that the rate control algorithm in the above video coding method is implemented under the HPM9.1 configuration. Since AVS3 does not currently include a rate control algorithm, this time the results are compared with the fixed QP. First, FixedQP encoding is performed on the HPM9.1 reference software at four QP points (27, 32, 38, 45) to obtain the bit rate of each sequence at each QP, and this bit rate is used as the target bit rate of the corresponding sequence. The final performance is shown in the following table (1). The experimental results show that the average error rate of the proposed code control model is only 0.44%. Compared with FixedQP, the BDrate of the three YUV channels is 0.35%, -1.95% and 2.45% respectively (positive values indicate performance degradation, negative values indicate performance improvement), and the weighted BDrate is -0.5% (the weights of the three channels are 4:1:1). That is, compared with traditional encoding methods, the above video encoding method has lower average error and better effect.

[0136]

[0137] Table (1);

[0138] It should be noted that for the aforementioned method embodiments, for simplicity of description, they are all expressed as a series of action combinations. However, those skilled in the art should be aware that the present invention is not limited by the order of the actions described, because according to the present invention, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in this specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the present invention.

[0139] According to another aspect of the embodiments of the present invention, a video encoding device for implementing the above-mentioned video encoding method is also provided. Figure 6 As shown, the device includes:

[0140] A determination unit 602 is configured to determine a video frame to be encoded from a video stream to be played;

[0141] an acquiring unit 604 configured to acquire, upon acquiring an average SATD value of the sum of the residual absolute values of each pixel in the to-be-encoded video frame, a target quantization parameter based on the average SATD value and a target number of bits corresponding to a maximum bit rate of the to-be-encoded video frame, wherein the average SATD value is used to indicate content complexity of the to-be-encoded video frame, and the target quantization parameter is used to indicate an output bit rate of the to-be-encoded video frame;

[0142] The encoding unit 606 is configured to encode the video frame to be encoded according to the output bit rate indicated by the target quantization parameter.

[0143] Optionally, in this embodiment, the video encoding device can be applied to, but not limited to, full I-frame rate control for the AVS3 video coding standard. It can also be applied to, but not limited to, video encoders, including software encoders and hardware encoders, to achieve stability and optimal rate-distortion performance for video transmission tasks under limited bandwidth constraints. Specifically, the quantization parameter (QP) of the I-frame is determined based on the SATDpp and target bits of the video frame, and on this basis, the output bit rate indicated by the quantization parameter of the I-frame is proposed to encode the video frame to be encoded, wherein the quantization parameter can be used to determine, but is not limited to, the quantization step size.

[0144] Optionally, in this embodiment, the above-mentioned video encoding device can also be applied to application scenarios such as video playback applications, video sharing applications or video conversation applications, but is not limited to it. The videos transmitted in the above-mentioned application scenarios may include but are not limited to: long videos and short videos. For example, long videos can be played for a long time (for example, the playing time is greater than 10 minutes) of the broadcasting series, or the pictures displayed in a long video conversation. Short videos can be voice messages for interaction between two or more parties, or videos with a short playing time (for example, the playing time is less than or equal to 30 seconds) for display on a sharing platform. The above is only an example. The video encoding device provided in this embodiment can be applied to, but is not limited to, the playback device for playing videos in the above-mentioned application scenarios. After obtaining the video frames to be encoded from the video stream to be played, the target quantization parameters of each video frame to be encoded are determined based on the average SATD value with low complexity to avoid the generation of prediction errors and ensure the accuracy of video encoding. Optionally, it is also possible to, but not limited to, perform bit allocation and QP determination on the three YUV channels separately.

[0145] Optionally, in this embodiment, in the video transmission task, the limited channel bandwidth is difficult to meet the growing demand for video resolution and variety. In order to stably and efficiently transmit the video bit stream, it is necessary to perform rate control at the encoding end. The rate control can be, but is not limited to, by selecting a series of quantization parameters (including frame level and block level), so as to achieve the stability of the video transmission task and the optimal rate-distortion performance under the condition of limited bandwidth constraints. The goal of the rate control work is equivalent to obtaining the quantization parameters through bit allocation and estimation of model parameters under the constraint of limited bandwidth, so as to minimize the overall distortion. Specifically, the video frame to be encoded is encoded according to the output rate indicated by the obtained quantization parameters.

[0146] Optionally, in this embodiment, the video frame to be encoded can be a video frame captured in real time, or a video frame corresponding to a stored video. The video frame to be encoded can be an input video frame in an input video frame sequence; the video frame to be encoded can also be a video frame obtained by processing an input video frame in an input video frame sequence according to a corresponding processing method, wherein, after the encoding end processes the input video frame according to the corresponding processing method, the resolution of the video frame to be encoded is smaller than the resolution of the original input video frame. For example, the input video frame can be downsampled according to a corresponding sampling ratio to obtain the video frame to be encoded.

[0147] Specifically, the encoding end can determine the processing method of the input video frame, and process the input video frame according to the processing method of the input video frame to obtain the video frame to be encoded. The processing methods include downsampling processing method and full-resolution processing method. The downsampling processing method refers to downsampling the input video frame to obtain the video frame to be encoded, and then encoding the obtained video frame to be encoded. The downsampling method in the downsampling processing method can be customized according to needs, including vertical downsampling, horizontal downsampling, vertical and horizontal downsampling, and downsampling can be performed using direct averaging, filters, bicubic interpolation, bilinear interpolation and other algorithms. The full-resolution processing method refers to directly using the input video frame as the video frame to be encoded, and directly encoding the video frame to be encoded based on the original resolution of the input video frame.

[0148] Optionally, in this embodiment, the device for determining the processing method corresponding to the input video frame can be configured based on actual needs. For example, processing parameters corresponding to the input video frame can be obtained, and the corresponding processing method can be determined based on the processing parameters. The processing parameters are parameters used to determine the processing method, and the specific processing method used can be configured based on actual needs. For example, the processing parameters may include current encoding information and / or image features corresponding to the input video frame.

[0149] Optionally, in this embodiment, the encoding end can obtain the processing method corresponding to the input video frame based on the current encoding information corresponding to the input video frame and at least one of the image feature information. The current encoding information refers to the video compression parameter information obtained when the video is encoded, such as one or more of the frame type, motion vector, quantization parameter, video source, bit rate, frame rate and resolution. Image feature information refers to information related to the image content, including one or more of the image motion information and image texture information, such as edges. The current encoding information and image feature information reflect the scene, detail complexity or motion intensity corresponding to the video frame. For example, the motion scene can be judged by one or more of the motion vector, quantization parameter or bit rate. A large quantization parameter generally indicates an intense motion, and a large motion vector indicates that the image scene is a large motion scene.

[0150] The average SATD value of the sum of the absolute residual values of each pixel in the video frame to be encoded is used as the representation of the image feature information. Since the average SATD value has low complexity, it can be obtained without encoding. Then, the average SATD value of the sum of the absolute residual values of each pixel in the video frame to be encoded is used to determine the texture complexity of the video frame to be encoded, eliminating the prediction error caused by the chicken-and-egg paradox problem and improving the accuracy of subsequent encoding-related calculations.

[0151] Optionally, in this embodiment, the target number of bits may be, but is not limited to, the expected number of bits obtained by calculating based on the maximum bit rate of the video frame to be encoded, wherein the bit rate (Bitrate) may be, but is not limited to, also known as the bit rate, represented by the letter r or R, indicating the rate of the bit stream after video encoding, and commonly used units are bps, Kbps, Mbps, and Gbps.

[0152] Optionally, in this embodiment, the sum of the absolute residual values of each pixel in the video frame to be encoded may be, but is not limited to, the sum of the absolute values of the residual absolute values after Hadamard transformation (Sum of Absolute Transformed Difference, SATD), which is used to represent the complexity of the video frame and the coding block content. The average SATD value may be, but is not limited to, the average SATD value per pixel, i.e., SATD per pixel.

[0153] It should be noted that, a video frame to be encoded is determined from a video stream to be played; when the average SATD value of the sum of the absolute values of the residuals of each pixel in the video frame to be encoded is obtained, a target quantization parameter is obtained based on the average SATD value and the target number of bits corresponding to the maximum bit rate of the video frame to be encoded, wherein the average SATD value is used to indicate the content complexity of the video frame to be encoded, and the target quantization parameter is used to indicate the output bit rate of the video frame to be encoded; the video frame to be encoded is encoded according to the output bit rate indicated by the target quantization parameter.

[0154] For specific embodiments, reference may be made to the examples shown in the above video encoding method, which will not be described in detail in this example.

[0155] Through the embodiments provided by the present application, a video frame to be encoded is determined from a video stream to be played; when the average SATD value of the sum of the absolute values of the residuals of each pixel in the video frame to be encoded is obtained, a target quantization parameter is obtained based on the average SATD value and the target number of bits corresponding to the maximum bit rate of the video frame to be encoded, wherein the average SATD value is used to indicate the content complexity of the video frame to be encoded, and the target quantization parameter is used to indicate the output bit rate of the video frame to be encoded; the video frame to be encoded is encoded according to the output bit rate indicated by the target quantization parameter, and by using the average SATD value that can be obtained without encoding as a measure of the content complexity of the video frame to be encoded, the chicken-and-egg paradox problem existing in the prior art is overcome, the prediction error is eliminated, and the technical purpose of improving the accuracy of the output bit rate indicated by the target quantization parameter is achieved, thereby achieving the technical effect of improving the encoding accuracy of the video.

[0156] As an alternative, for example Figure 7 As shown, the acquisition unit 604 includes:

[0157] A calculation module 702 is configured to calculate a target number of bits and an average SATD value to obtain a frame-level quantization parameter, wherein the frame-level quantization parameter is used to indicate a frame-level output bit rate of a video frame to be encoded;

[0158] The first determining module 704 is configured to determine the frame-level quantization parameter as the target quantization parameter; or

[0159] an acquisition module 706 for acquiring a block-level quantization parameter based on the frame-level quantization parameter obtained by calculating the target number of bits and the average SATD value, wherein the block-level quantization parameter is used to indicate a block-level output bit rate of the video frame to be encoded;

[0160] The second determining module 708 is configured to determine the block-level output bit rate as a target quantization parameter.

[0161] For specific embodiments, reference may be made to the examples shown in the above video encoding method, which will not be described in detail in this example.

[0162] As an optional solution, the calculation module 702 includes:

[0163] a first calculation submodule, configured to calculate a target number of bits and an average SATD value using a target mathematical formula to obtain a frame-level quantization parameter when the video frame to be encoded is a first video frame, wherein the target mathematical formula is used to represent a mathematical relationship between the average SATD value, the target number of bits, and the frame-level quantization parameter; or

[0164] an updating submodule, configured to update a calculation parameter in a target mathematical formula based on the number of bits consumed by the first video frame and a target number of bits when the video frame to be encoded is a second video frame, wherein the second video frame is a next video frame of the first video frame;

[0165] The second calculation submodule is used to calculate the target number of bits and the average SATD value using the updated target mathematical formula to obtain the frame-level quantization parameter.

[0166] For specific embodiments, reference may be made to the examples shown in the above video encoding method, which will not be described in detail in this example.

[0167] As an optional solution, the acquisition module 706 includes:

[0168] The acquisition submodule is used to update the frame-level quantization parameter using the reference quantization parameter to obtain the block-level quantization parameter when the reference quantization parameter of the video frame to be encoded is obtained, wherein the reference quantization parameter is used to adjust the bit rate error of the frame-level quantization parameter.

[0169] For specific embodiments, reference may be made to the examples shown in the above video encoding method, which will not be described in detail in this example.

[0170] As an optional solution, get the submodules, including:

[0171] The first determining subunit is configured to determine the frame-level quantization parameter as the block-level quantization parameter when the video frame to be encoded is the third video frame; or

[0172] a first updating subunit, configured to update a reference quantization parameter based on the number of bits consumed by the third video frame when the video frame to be encoded is the fourth video frame;

[0173] A second updating subunit, configured to update the frame-level quantization parameter using the updated reference quantization parameter;

[0174] The second determining subunit is configured to determine the updated frame-level quantization parameter as the block-level quantization parameter.

[0175] For specific embodiments, reference may be made to the examples shown in the above video encoding method, which will not be described in detail in this example.

[0176] As an optional solution, the first determining subunit includes at least one of the following:

[0177] A first sub-updating module is configured to update the reference quantization parameter to a first reference quantization parameter when the number of bits consumed by the third video frame does not reach a target number of bits, wherein the first reference quantization parameter is used to indicate a frame-level quantization parameter whose average SATD value is greater than or equal to a complexity threshold to be lowered;

[0178] The second sub-update module is used to update the reference quantization parameter to a second reference quantization parameter when the number of bits consumed by the third video frame does not reach the target number of bits, wherein the second reference quantization parameter is used to indicate an increase in the frame-level quantization parameter whose average SATD value is less than the complexity threshold.

[0179] For specific embodiments, reference may be made to the examples shown in the above video encoding method, which will not be described in detail in this example.

[0180] As an alternative, for example Figure 8 As shown, the acquisition unit 604 includes:

[0181] A third determining module 802 is configured to determine an upper limit quantization parameter and a lower limit quantization parameter of a target quantization parameter upon obtaining a previous frame-level quantization parameter corresponding to a previous video frame of the video frame to be encoded, wherein the upper limit quantization parameter is the sum of the previous frame-level quantization parameter and a preset threshold, and the lower limit quantization parameter is the difference between the previous frame-level quantization parameter and the preset threshold, and the preset threshold is used to constrain the adjustment range of the quantization parameter;

[0182] A fourth determining module 804 is configured to use the upper limit quantization parameter as the target quantization parameter when the target quantization parameter is greater than or equal to the upper limit quantization parameter;

[0183] The fifth determining module 806 is configured to use the lower limit quantization parameter as the target quantization parameter when the target quantization parameter is less than or equal to the lower limit quantization parameter.

[0184] For specific embodiments, reference may be made to the examples shown in the above video encoding method, which will not be described in detail in this example.

[0185] According to another aspect of the embodiment of the present invention, an electronic device for implementing the above-mentioned video encoding method is also provided. Figure 9 As shown, the electronic device includes a memory 902 and a processor 904. The memory 902 stores a computer program, and the processor 904 is configured to execute the steps in any of the above method embodiments through the computer program.

[0186] Optionally, in this embodiment, the electronic device may be located in at least one network device among a plurality of network devices of a computer network.

[0187] Optionally, in this embodiment, the processor may be configured to execute the following steps through a computer program:

[0188] S1, determining a video frame to be encoded from a video stream to be played;

[0189] S2, when an average SATD value of the sum of the residual absolute values of each pixel in the video frame to be encoded is obtained, obtaining a target quantization parameter according to the average SATD value and a target number of bits corresponding to the maximum bit rate of the video frame to be encoded, wherein the average SATD value is used to indicate the content complexity of the video frame to be encoded, and the target quantization parameter is used to indicate the output bit rate of the video frame to be encoded;

[0190] S3, encoding the video frame to be encoded according to the output bit rate indicated by the target quantization parameter.

[0191] Alternatively, those skilled in the art will appreciate that Figure 9 The structure shown is for illustration only, and the electronic device may also be a smart phone (such as an Android phone, an iOS phone, etc.), a tablet computer, a PDA, a mobile Internet device (MID), a PAD, or other terminal devices. Figure 9 It does not limit the structure of the above electronic device. For example, the electronic device may also include Figure 9 More or fewer components (such as network interfaces, etc.) as shown in, or with Figure 9 Different configurations shown.

[0192] Among them, the memory 902 can be used to store software programs and modules, such as program instructions / modules corresponding to the video encoding method and device in the embodiment of the present invention. The processor 904 executes various functional applications and data processing by running the software programs and modules stored in the memory 902, that is, realizing the above-mentioned video encoding method. The memory 902 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 902 may further include a memory remotely located relative to the processor 904, and these remote memories may be connected to the terminal via a network. Examples of the above-mentioned networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof. Among them, the memory 902 can be used specifically, but not limited to, to store information such as video frames to be encoded, target quantization parameters, and output bit rates. As an example, if Figure 9 As shown, the memory 902 may include, but is not limited to, the determining unit 602, the acquiring unit 604, and the encoding unit 606 in the video encoding apparatus. Furthermore, it may also include, but is not limited to, other modules and units in the video encoding apparatus, which will not be described in detail in this example.

[0193] Optionally, the transmission device 906 is used to receive or send data via a network. Specific examples of the network may include wired networks and wireless networks. In one embodiment, the transmission device 906 includes a network interface controller (NIC), which can be connected to other network devices and a router via a network cable to communicate with the Internet or a local area network. In one embodiment, the transmission device 906 is a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.

[0194] In addition, the electronic device further includes: a display 908 for displaying information such as the video frame to be encoded, target quantization parameter, and output bit rate; and a connection bus 910 for connecting various module components in the electronic device.

[0195] In other embodiments, the terminal device or server may be a node in a distributed system, wherein the distributed system may be a blockchain system, and the blockchain system may be a distributed system formed by connecting multiple nodes via network communication. The nodes may form a peer-to-peer (P2P) network, and any computing device, such as a server, terminal, or other electronic device, may become a node in the blockchain system by joining the peer-to-peer network.

[0196] According to one aspect of the present application, a computer program product or computer program is provided, comprising computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the above-described video encoding method, wherein the computer program is configured to perform the steps of any of the above-described method embodiments when executed.

[0197] Optionally, in this embodiment, the computer-readable storage medium may be configured to store a computer program for performing the following steps:

[0198] S1, determining a video frame to be encoded from a video stream to be played;

[0199] S2, when an average SATD value of the sum of the residual absolute values of each pixel in the video frame to be encoded is obtained, obtaining a target quantization parameter according to the average SATD value and a target number of bits corresponding to the maximum bit rate of the video frame to be encoded, wherein the average SATD value is used to indicate the content complexity of the video frame to be encoded, and the target quantization parameter is used to indicate the output bit rate of the video frame to be encoded;

[0200] S3, encoding the video frame to be encoded according to the output bit rate indicated by the target quantization parameter.

[0201] Optionally, in this embodiment, a person of ordinary skill in the art may understand that all or part of the steps in the various methods of the above embodiments may be completed by instructing the hardware related to the terminal device through a program, and the program may be stored in a computer-readable storage medium, which may include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.

[0202] The serial numbers of the above embodiments of the present invention are for description only and do not represent the advantages or disadvantages of the embodiments.

[0203] If the integrated units in the above embodiments are implemented in the form of software functional units and sold or used as independent products, they can be stored in the above-mentioned computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the existing technology, or all or part of the technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes a number of instructions for causing one or more computer devices (such as personal computers, servers, or network devices) to execute all or part of the steps of the methods described in various embodiments of the present invention.

[0204] In the above embodiments of the present invention, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0205] In the several embodiments provided in this application, it should be understood that the disclosed client can be implemented in other ways. Among them, the device embodiments described above are merely illustrative. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, and can be electrical or other forms.

[0206] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0207] In addition, the functional units in the various embodiments of the present invention may be integrated into a single processing unit, each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0208] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be regarded as within the scope of protection of the present invention.

Claims

1. A video encoding method, characterized in that: include: Determine the video frame to be encoded from the video stream to be played; Upon obtaining an average SATD value of the sum of the residual absolute values of each pixel in the to-be-encoded video frame, obtaining a frame-level quantization parameter by calculating a target number of bits corresponding to a maximum bit rate of the to-be-encoded video frame and the average SATD value, and obtaining a reference quantization parameter of the to-be-encoded video frame, and when the to-be-encoded video frame is not a first frame, determining the frame-level quantization parameter as a block-level quantization parameter, wherein the block-level quantization parameter is used to indicate a block-level output bit rate of the to-be-encoded video frame, and the reference quantization parameter is used to adjust a bit rate error of the frame-level quantization parameter; When the average SATD value is obtained, the frame-level quantization parameter is obtained by calculating the target number of bits and the average SATD value, and the reference quantization parameter is obtained, and when the video frame to be encoded is the first frame, updating the reference quantization parameter based on the number of bits consumed by the video frame to be encoded; Updating the frame-level quantization parameter using the updated reference quantization parameter; Determining the updated frame-level quantization parameter as the block-level quantization parameter; Determining the block-level output bit rate as a target quantization parameter; The to-be-encoded video frame is encoded according to the output bit rate indicated by the target quantization parameter.

2. The method according to claim 1, characterized in that Before determining the frame-level quantization parameter as the block-level quantization parameter, the method further includes: When the video frame to be encoded is the first video frame, calculating the target number of bits and the average SATD value using a target mathematical formula to obtain the frame-level quantization parameter, wherein the target mathematical formula is used to represent a mathematical relationship between the average SATD value, the target number of bits, and the frame-level quantization parameter; or When the video frame to be encoded is a second video frame, updating calculation parameters in the target mathematical formula based on the number of bits consumed by the first video frame and the target number of bits, wherein the second video frame is a next video frame of the first video frame; The target number of bits and the average SATD value are calculated using the updated target mathematical formula to obtain the frame-level quantization parameter.

3. The method according to claim 1, characterized in that The updating of the reference quantization parameter based on the number of bits consumed by the to-be-encoded video frame includes at least one of the following: When the number of bits consumed by the to-be-encoded video frame does not reach the target number of bits, updating the reference quantization parameter to a first reference quantization parameter, wherein the first reference quantization parameter is used to indicate to lower the frame-level quantization parameter whose average SATD value is greater than or equal to a complexity threshold; When the number of bits consumed by the video frame to be encoded is greater than the target number of bits, the reference quantization parameter is updated to a second reference quantization parameter, wherein the second reference quantization parameter is used to indicate to increase the frame-level quantization parameter whose average SATD value is less than the complexity threshold.

4. The method according to any one of claims 1 to 3, characterized in that The obtaining of a target quantization parameter according to the average SATD value and a target number of bits corresponding to a maximum bit rate of the video frame to be encoded includes: determining an upper limit quantization parameter and a lower limit quantization parameter of the target quantization parameter when a previous frame-level quantization parameter corresponding to a previous video frame of the video frame to be encoded is obtained, wherein the upper limit quantization parameter is the sum of the previous frame-level quantization parameter and a preset threshold, and the lower limit quantization parameter is the difference between the previous frame-level quantization parameter and the preset threshold, and the preset threshold is used to constrain the adjustment range of the quantization parameter; When the target quantization parameter is greater than or equal to the upper limit quantization parameter, using the upper limit quantization parameter as the target quantization parameter; When the target quantization parameter is less than or equal to the lower limit quantization parameter, the lower limit quantization parameter is used as the target quantization parameter.

5. A video encoding device, characterized in that: include: A determination unit, configured to determine a video frame to be encoded from a video stream to be played; an acquisition unit, configured to, upon obtaining an average SATD value of the sum of the residual absolute values of each pixel in the to-be-encoded video frame, obtaining a frame-level quantization parameter by calculating a target number of bits corresponding to a maximum bit rate of the to-be-encoded video frame and the average SATD value, obtaining a reference quantization parameter for the to-be-encoded video frame, and, if the to-be-encoded video frame is not a first frame, determine the frame-level quantization parameter as a block-level quantization parameter, wherein the block-level quantization parameter is used to indicate a block-level output bit rate of the to-be-encoded video frame, and the reference quantization parameter is used to adjust a bit rate error of the frame-level quantization parameter; upon obtaining the average SATD value, obtaining the frame-level quantization parameter by calculating the target number of bits and the average SATD value, obtaining the reference quantization parameter, and if the to-be-encoded video frame is a first frame, update the reference quantization parameter based on the number of bits consumed by the to-be-encoded video frame; update the frame-level quantization parameter using the updated reference quantization parameter; determine the updated frame-level quantization parameter as the block-level quantization parameter; and determine the block-level output bit rate as the target quantization parameter; The encoding unit is configured to encode the to-be-encoded video frame according to the output bit rate indicated by the target quantization parameter.

6. A computer-readable storage medium, characterized in that The computer-readable storage medium includes a stored program, wherein the program executes the method described in any one of claims 1 to 4 when executed.

7. An electronic device comprising a memory and a processor, characterized in that: A computer program is stored in the memory, and the processor is configured to execute the method according to any one of claims 1 to 4 through the computer program.

Citation Information

Patent Citations

  • Optimization method and system for code rate control of video monitor in 3G network

    CN103841418A

  • Code rate control method and device based on layered B frame and electronic equipment

    CN109862359A