Video encoding method and apparatus, and electronic device and storage medium
By enhancing the content of video frames and optimizing the encoding parameters, the quality degradation caused by video encoding in existing technologies is solved, and the image quality and visual effects of the encoded video are improved.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- TENCENT TECHNOLOGY (SHENZHEN) CO LTD
- Filing Date
- 2025-09-09
- Publication Date
- 2026-05-21
AI Technical Summary
Existing video encoding technologies lose many details of video frames during compression, resulting in a decline in video quality and visual effects, which affects user experience.
By enhancing the content of the video frame to be encoded, detecting the enhancement information of the enhanced video frame, calculating the encoding parameters based on this information, encoding the video frame, and generating encoded data.
It improves the image quality and visual effects of encoded video frames, reduces the loss of video frame details, and improves the overall quality of the encoded video.
Smart Images

Figure CN2025120094_21052026_PF_FP_ABST
Abstract
Description
Video encoding methods and apparatus, electronic devices, storage media
[0001] Related applications
[0002] This application claims priority to Chinese patent application filed on November 13, 2024, with application number 202411629627.3, entitled "Video Coding Method and Apparatus, Electronic Device, Storage Medium", the entire contents of which are incorporated herein by reference. Technical Field
[0003] This application relates to the field of computers, and more specifically, to a video encoding method and apparatus, electronic device, storage medium, and program product. Background Technology
[0004] With the rapid development of the video industry, video is being used in an increasing number of scenarios, such as short videos, e-commerce live streaming, live show broadcasts, game streaming, and real-time cloud rendering. This places increasingly higher demands on video clarity and frame rate. To facilitate video storage, playback, and transmission, video encoding is necessary.
[0005] In related technologies, video frames are compressed during the video encoding process, resulting in the loss of many details in the video frames, which reduces the quality of the video frames and the visual effect of the video, thus affecting the user experience. Summary of the Invention
[0006] Embodiments of this application provide a video encoding method and apparatus, an electronic device, a storage medium, and a program product.
[0007] According to one aspect of the embodiments of this application, a video encoding method is provided, the method comprising:
[0008] Content enhancement is performed on the original video frames contained in the video to be encoded to obtain enhanced video frames;
[0009] The enhanced video frame is detected to obtain its content enhancement information, and the encoding parameters of the enhanced video frame are calculated based on the content enhancement information; and
[0010] The enhanced video frame is encoded according to the enhanced video frame encoding parameters to obtain the encoded data of the enhanced video frame, and the encoded data of the video to be encoded is generated based on the encoded data of the enhanced video frame.
[0011] According to one aspect of the embodiments of this application, a video encoding apparatus is provided, the apparatus comprising:
[0012] The enhancement module is configured to enhance the content of the original video frames contained in the video to be encoded, thereby obtaining enhanced video frames;
[0013] The processing module is configured to detect the enhanced video frame, obtain content enhancement information of the enhanced video frame, and calculate the encoding parameters of the enhanced video frame based on the content enhancement information; and
[0014] The encoding module is configured to encode the enhanced video frame according to the enhanced video frame encoding parameters to obtain the encoded data of the enhanced video frame, and generate the encoded data of the video to be encoded based on the encoded data of the enhanced video frame.
[0015] According to one aspect of the embodiments of this application, an electronic device is provided, comprising:
[0016] At least one processor;
[0017] A storage device for storing at least one computer program, which, when executed by the at least one processor, causes the electronic device to implement the video encoding method as described above.
[0018] According to one aspect of the embodiments of this application, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor of an electronic device, causes the electronic device to implement the video encoding method as described above.
[0019] According to one aspect of the embodiments of this application, a computer program product is provided, including a computer program that, when executed by a processor, implements the video encoding method as described above.
[0020] According to one aspect of the embodiments of this application, a method for processing a video stream is provided, the video stream being generated according to the video encoding method described above.
[0021] According to one aspect of the embodiments of this application, a computer-readable storage medium is provided having computer instructions stored thereon, which, when executed by at least one processor, implement the video encoding method as described above to generate and store a video stream.
[0022] Details of one or more embodiments of this application are set forth in the following drawings and description. Other features, objects, and advantages of this application will become apparent from the specification, drawings, and claims. Attached Figure Description
[0023] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only embodiments of this application. For those skilled in the art, other drawings can be obtained based on the published drawings without creative effort.
[0024] Figure 1 is a schematic diagram illustrating an implementation environment of an exemplary embodiment of this application;
[0025] Figure 2 is a flowchart illustrating a video encoding method in an exemplary embodiment of this application;
[0026] Figure 3 is a schematic diagram of the model structure shown in an exemplary embodiment of this application;
[0027] Figure 4 is a schematic diagram illustrating a video encoding process in an exemplary embodiment of this application;
[0028] Figure 5 is a schematic diagram illustrating the relationship between CTU, PU, and TU in an exemplary embodiment of this application;
[0029] Figure 6 is a flowchart illustrating a video encoding method in another exemplary embodiment of this application;
[0030] Figure 7 is a flowchart illustrating a video encoding method in another exemplary embodiment of this application;
[0031] Figure 8 is a flowchart illustrating a video encoding method in another exemplary embodiment of this application;
[0032] Figure 9A is a flowchart illustrating a video encoding method in another exemplary embodiment of this application;
[0033] Figure 9B is a schematic diagram of a video frame illustrating an exemplary embodiment of this application;
[0034] Figure 10 is a flowchart illustrating a video encoding method in another exemplary embodiment of this application;
[0035] Figure 11 is a flowchart illustrating a video encoding method in another exemplary embodiment of this application;
[0036] Figure 12A is a schematic diagram illustrating the rate control process in an exemplary embodiment of this application;
[0037] Figure 12B is a flowchart illustrating the calculation of quantization parameters in an exemplary embodiment of this application;
[0038] Figure 13 is a flowchart illustrating a video encoding method in another exemplary embodiment of this application;
[0039] Figure 14 is a schematic diagram illustrating a video encoding process in another exemplary embodiment of this application;
[0040] Figure 15 is a schematic diagram of a video encoding apparatus illustrating an exemplary embodiment of this application;
[0041] Figure 16 shows a schematic diagram of the structure of a computer system suitable for implementing the electronic device of the present application. Detailed Implementation
[0042] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0043] Exemplary embodiments will now be described in a more comprehensive manner with reference to the accompanying drawings. However, the exemplary embodiments can be implemented in various forms and should not be construed as limited to these examples; rather, these embodiments are provided so that this application will be more comprehensive and complete, and will fully convey the concept of the exemplary embodiments to those skilled in the art.
[0044] Furthermore, the features, structures, or characteristics described in this application can be combined in any suitable manner in one or more embodiments. Numerous specific details are provided in the following description to provide a full understanding of the embodiments of this application. However, those skilled in the art will recognize that when implementing the technical solutions of this application, not all the detailed features in the embodiments may be used, one or more specific details may be omitted, or other methods, elements, devices, steps, etc., may be employed.
[0045] In the embodiments of this application, the terms "module" or "unit" refer to a computer program or part of a computer program that has a predetermined function and works with other related parts to achieve a predetermined goal, and can be implemented wholly or partially using software, hardware (such as processing circuitry or memory), or a combination thereof. Similarly, a processor (or at least two processors or memory) can be used to implement one or at least two modules or units. Furthermore, each module or unit can be part of an overall module or unit that includes the functionality of that module or unit.
[0046] The block diagrams shown in the accompanying drawings are merely functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software, in one or at least two hardware modules or integrated circuits, or in different network and / or processor devices and / or microcontroller devices.
[0047] The flowcharts shown in the accompanying drawings are merely illustrative and do not necessarily include all content and operations / steps, nor do they necessarily have to be performed in the described order. For example, some operations / steps can be broken down, while others can be combined or partially combined; therefore, the actual execution order may change depending on the specific circumstances.
[0048] It should also be noted that "multiple" as mentioned in this application refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone. The character " / " generally indicates that the preceding and following related objects have an "or" relationship.
[0049] The technical solutions of the embodiments of this application are described in detail below:
[0050] In related technologies, encoding video significantly reduces its quality. Therefore, embodiments of this application provide a video encoding method and apparatus, electronic device, storage medium, and program product that can improve the quality and visual effects of the encoded video.
[0051] Please refer to Figure 1, which is a schematic diagram of an implementation environment related to this application. This implementation environment includes a terminal device 110 and a server 120. The terminal device 110 and the server 120 communicate with each other via a wired or wireless network. The terminal device 110 can upload its own data to the server 120 and can also retrieve data from the server 120.
[0052] Among them, terminal device 110 may include, but is not limited to, mobile phones, tablets, laptops, computers, voice interaction devices, home appliances, vehicle terminals, aircraft, remote driving terminals, etc.; server 120 may be an independent physical server, or a server cluster or distributed system consisting of at least two physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. This document does not restrict the specific form of terminal devices and servers.
[0053] It should be noted that the number of terminal devices 110 and servers 120 in Figure 1 is merely illustrative. Depending on actual needs, there can be any number of terminal devices 110 and servers 120.
[0054] In an exemplary embodiment, the video encoding method provided in the embodiments of this application can be executed by a terminal device 110. Exemplarily, the terminal device 110 can first perform content enhancement on the original video frames contained in the video to be encoded to obtain enhanced video frames. It can then detect the enhanced video frames to obtain content enhancement information, calculate the encoding parameters of the enhanced video frames based on the content enhancement information, and encode the enhanced video frames according to the encoding parameters to obtain encoded data of the enhanced video frames. Based on the encoded data of the enhanced video frames, it can generate encoded data for the video to be encoded. On the one hand, performing content enhancement on the video frames before encoding can improve the image quality of the encoded video frames. On the other hand, the encoding parameters of the video frames are calculated based on the content enhancement information, thereby deeply integrating the content enhancement process with the encoding process, reducing the weakening effect of the encoding process on the enhancement effect, improving the enhancement effect of the encoded video frames, reducing the loss of video frame details, further improving the quality of the encoded video frames, and improving the quality and visual effect of the encoded video.
[0055] In another exemplary embodiment, server 120 may have functions similar to terminal device 110 to execute the video encoding method provided in this application embodiment. Exemplarily, server 120 may first perform content enhancement on the original video frames contained in the video to be encoded to obtain enhanced video frames, then detect the enhanced video frames to obtain content enhancement information, and calculate the encoding parameters of the enhanced video frames based on the content enhancement information. The enhanced video frames are then encoded according to the encoding parameters to obtain encoded data of the enhanced video frames, and encoded data of the video to be encoded is generated based on the encoded data of the enhanced video frames.
[0056] In another exemplary embodiment, the terminal device 110 and the server 120 may also jointly execute the video encoding method provided in the embodiments of this application. For example, the terminal device 110 may acquire a video to be encoded and send it to the server 120; the server 120 may perform content enhancement on the original video frames contained in the video to be encoded to obtain enhanced video frames, detect the enhanced video frames to obtain content enhancement information of the enhanced video frames, calculate the encoding parameters of the enhanced video frames based on the content enhancement information, and encode the enhanced video frames according to the encoding parameters to obtain encoded data of the enhanced video frames, so as to generate encoded data of the video to be encoded based on the encoded data of the enhanced video frames.
[0057] The video encoding method in this application can be applied to video encoding scenarios in various applications. The embodiments of this application involve user-related data such as the video to be encoded. When the method of this application is applied to specific products or technologies, user permission or consent is always obtained, and the extraction, use, and processing of related data comply with local security standards and local laws and regulations.
[0058] Referring to Figure 2, which is a flowchart illustrating a video encoding method in an exemplary embodiment of this application, the method can be applied to the implementation environment shown in Figure 1. It can be executed by the terminal device 110 in the implementation environment shown in Figure 1, or by the server 120 in the implementation environment shown in Figure 1, or by both the terminal device 110 and the server 120 in the implementation environment shown in Figure 1.
[0059] As shown in Figure 2, in an exemplary embodiment, the video encoding method may include steps S210-S230, which are described in detail below:
[0060] Step S210: Perform content enhancement on the original video frames contained in the video to be encoded to obtain enhanced video frames.
[0061] It should be noted that "video to be encoded" refers to any video that needs encoding, such as videos that need encoding for transmission, videos that need transcoding, or videos that need compression. Based on the video's creator, the type of video to be encoded includes, but is not limited to, User-Generated Content (UGC) and Professionally-Generated Content (PGC). UGC refers to content created by ordinary users, while PGC refers to content created by professional organizations (e.g., broadcasting professionals). Based on the video's length, the type of video to be encoded includes, but is not limited to, short videos and long videos. Short videos are those lasting a few minutes or even tens of seconds, while long videos are those lasting tens of minutes or even hours. Based on real-time nature, the type of video to be encoded includes, but is not limited to, live videos and non-live videos. Live videos are those recorded and broadcast in real time, while non-live videos are those created, stored, and available for viewing at any time.
[0062] By parsing the video to be encoded, we can obtain the video frames contained in the video, that is, the original video frames. By enhancing the content of each original video frame, we can obtain the enhanced video frames.
[0063] Content enhancement is a preprocessing step used to improve video quality, clarity, and overall viewing experience. Specific methods include, but are not limited to, super-resolution processing, noise reduction, frame interpolation, color enhancement, deblurring, font adjustment, decompression, artifact removal, subtitle addition, watermarking, face recognition, comprehensive enhancement, old film restoration, and image quality enhancement. Super-resolution is a technique for increasing the resolution of image or video frames, while frame interpolation is a technique that inserts extra frames between video frames using algorithms.
[0064] Content enhancement can be implemented in various ways, including but not limited to traditional CV (computer vision) models, deep learning models, and large-scale models. Traditional CV models primarily refer to computer vision techniques in the field of artificial intelligence that are not based on deep learning. Large-scale models refer to machine learning models with large-scale parameters and complex computational structures, including but not limited to Large Language Models (LLM), Diffusion Transformer (DIT) models, and large models combining LLM and DIT models. The DIT model is a diffusion model incorporating the Transformer architecture. Deep learning models are based on methods and algorithms using deep neural networks, including but not limited to Deep Neural Networks (DNN), Convolutional Neural Networks (CNN), Recurrent Neural Networks (RNN), Long Short-Term Memory (LSTM), Super-Resolution Convolutional Neural Networks (SRCNN), Generative Adversarial Networks (GAN), and Deep Reinforcement Learning (DRL). Super-resolution convolutional neural networks (CNNs) are deep neural network-based models used to improve the resolution of images or video frames. They consist of a patch extraction and representation module, a non-linear mapping module, and a reconstruction module. During super-resolution processing, the original video frames are first preprocessed to obtain low-resolution images, which are then input to the patch extraction and representation module for feature extraction. Finally, the images pass through the non-linear mapping and reconstruction modules, outputting a high-resolution image.
[0065] Taking histogram equalization in traditional CV models for color enhancement as an example, when enhancing the color of the original video frames, for each original video frame, it is converted from the RGB color space to the YCbCr color space. Since the Y channel represents luminance information, histogram equalization of the Y channel can effectively enhance the overall contrast of the image. The specific steps are as follows: Calculate the histogram H of the pixel values in the Y channel, where H[i] represents the number of pixels with a value of i, and the value of i ranges from 0 to 255. Calculate the cumulative distribution function (CDF). The value of j also ranges from 0 to 255. The cumulative distribution function is then normalized. CDF min This is the minimum value in the CDF, and N is the total number of pixels in the Y channel. Each pixel value in the Y channel is mapped according to the normalized cumulative distribution function to obtain the enhanced Y channel pixel value. Finally, the enhanced Y channel is merged with the original Cb and Cr channels and converted back to the RGB color space to obtain the color-enhanced video frame.
[0066] In an optional example, to improve the efficiency of super-resolution processing and the quality of the resulting high-resolution image, SRCNN is used to perform super-resolution processing on the original video frames. Referring to Figure 3, SRCNN can include a patch extraction and representation module, a non-linear mapping module, and a reconstruction module. During the super-resolution process, the original video frames can be preprocessed to obtain an image of the target size. This target-size image is then used as the low-resolution image to be improved and input into the patch extraction and representation module. The preprocessing method can be interpolation, for example, using a traditional CV model to interpolate the original video frames. Interpolation methods include, but are not limited to, bicubic interpolation and bilinear interpolation. Then, the patch extraction and representation module extracts features from the low-resolution image and outputs the corresponding feature map F1[Y]. The non-linear mapping module performs non-linear mapping on F1[Y] to obtain the feature map F2[Y]. The reconstruction module reconstructs the image based on F2[Y] and outputs the reconstructed image, i.e., the high-resolution image F[Y].
[0067] The block extraction and representation module can be a convolutional neural network, containing convolutional layers and activation layers. The activation function corresponding to the activation layer can be a rectified linear unit (ReLU). The mapping function corresponding to the block extraction and representation module can be as follows: F1[Y]=max(0,W1*Y+B1)
[0068] Where W1 and B1 represent the weight and bias parameters of the filters (i.e., the convolutional kernels of the convolutional layers) in the block extraction and representation modules, respectively, the max function corresponds to the ReLU function, and Y is the low-resolution image.
[0069] The nonlinear mapping module can be a convolutional neural network, containing convolutional layers and activation layers, and its corresponding mapping function can be as follows: F2[Y]=max(0,W2*F1[Y]+B2)
[0070] Where W2 and B2 represent the weight parameters and bias parameters of the filter (i.e., the convolution kernel of the convolutional layer) in the nonlinear mapping module, respectively.
[0071] The reconstruction module can be a convolutional neural network containing convolutional layers, and its corresponding mapping function can be as follows: F[Y]=W3*F2[Y]+B3
[0072] W3 and B3 represent the weight and bias parameters of the filters (i.e., the convolutional kernels of the convolutional layers) in the reconstruction module, respectively.
[0073] It should be noted that in other examples, the original video frames can also be directly input as low-resolution images into the block extraction and representation module for processing to obtain high-resolution images.
[0074] During the training of SRCNN, mean squared error (MSE) can be used as the loss function, and the model parameters of SRCNN can be updated based on the loss function using algorithms such as gradient descent. Optionally, to avoid boundary effects during training, the sample images can be left unpadded.
[0075] Step S220: Detect the enhanced video frame to obtain the content enhancement information of the enhanced video frame, and calculate the encoding parameters of the enhanced video frame based on the content enhancement information.
[0076] Content enhancement information refers to information related to content enhancement, including but not limited to the enhanced areas, the enhancement magnitude of the enhanced areas, the area type, importance, and level.
[0077] The encoding parameters for enhanced video frames refer to the encoding parameters used in the encoding process of enhanced video frames. The encoding process of video frames mainly includes the following steps:
[0078] Image segmentation: Dividing video frames into different unit blocks, such as macroblocks or coding tree units (CTUs), to facilitate subsequent processing.
[0079] Prediction: Predicting the pixel parameters of a unit block to obtain the predicted value of the unit block. Prediction types include intra-frame prediction and inter-frame prediction. Intra-frame prediction is a technique that uses the correlation between adjacent pixels in an image to eliminate spatial redundancy and improve compression efficiency. It mainly predicts the pixels to be encoded in the current frame based on the pixels already encoded in the current frame to reduce spatial redundancy. Intra-frame prediction includes different prediction modes, such as Planar mode, DC mode, and angle mode. Inter-frame prediction is a technique that uses the correlation (i.e., temporal correlation) between video frames to eliminate temporal redundancy and achieve image compression. It mainly finds a block from the already encoded frame to predict the pixels to be encoded in the current frame to reduce temporal redundancy. Since there is a strong correlation between video frames, using this characteristic for inter-frame coding can achieve a high compression ratio. Inter-frame prediction includes different prediction modes, such as forward prediction mode and bidirectional prediction mode. Forward prediction mode only refers to the preceding video frame, while bidirectional prediction mode can refer to the preceding and / or following video frames.
[0080] Let's take the three-step search method for motion estimation as an example. Motion estimation is a technique used in inter-frame prediction in video coding to find the block in already encoded frames that best matches the pixel block to be encoded in the current frame, thus determining the motion vector. The three-step search method is a specific implementation of motion estimation in inter-frame prediction of video coding. It is used to find the block in already encoded frames that best matches the pixel block to be encoded in the current frame, thus determining the motion vector, which will be used in subsequent inter-frame prediction and image compression processes.
[0081] In the three-step search method, the matching error is calculated in a specific area of the reference frame to gradually narrow the search range, and finally the best matching position of the current coding unit in the reference frame is found. The offset between this position and the center position of the current coding unit is the motion vector. When performing inter-frame prediction on a certain coding unit (CU) of the current frame, assume the reference frame is the previous frame. The specific steps are as follows: Step 1: Taking the center position of the current CU as the starting point, set a relatively large search step S1, and calculate the matching error (such as sum of absolute differences SAD) between each position in the nine-grid position (i.e., the up, down, left, right, and four diagonal positions) centered on this starting point in the reference frame, and select the position with the minimum matching error as the new center point. Step 2: Taking the new center point as the center, set a smaller search step S2 (S2 < S1), calculate the matching error again in the nine-grid positions around the new center point, and select the position with the minimum matching error as the new center point. Step 3: Taking the new center point obtained in the second step as the center, set an even smaller search step S3 (S3 < S2), calculate the matching error in the nine-grid positions around it, and finally select the position with the minimum matching error as the best matching position of the current CU in the reference frame. The offset between this best matching position and the center position of the current CU is the motion vector, and the predicted block is obtained from the reference frame using this motion vector to predict the pixel value of the current CU.
[0082] Transformation: Based on the predicted value and the original value of the unit block, calculate the residual parameters, and perform processing such as discrete cosine transform (DCT) on the residual parameters to further compress the data.
[0083] Quantization: Quantize the transformed data to convert continuous values into discrete values for easy coding.
[0084] Entropy coding: Perform entropy coding on the quantized data, such as Huffman coding or CABAC coding, to further reduce the data volume.
[0085] Bitstream format: Organize the encoded data into a bitstream according to a specific format for easy storage and transmission.
[0086] Exemplarily, as shown in FIG. 4, during the encoding process of the current frame F n (i.e., the nth video frame), F can be encoded first nThe partition is divided into CTUs, each CTU containing a Coding Unit (CU), and each CU containing a Predict Unit (PU) and a Transform Unit (TU). The relationship between CTUs, CUs, PUs, and TUs is shown in Figure 5. Then, intra-frame prediction or inter-frame prediction is used to predict the pixel parameters of the PUs, obtaining the predicted values. The predicted values are subtracted from the original values to obtain the residual parameters. The residual parameters are then subjected to DCT transformation and quantization to obtain residual coefficients. These residual coefficients are fed into the entropy coding module for encoding to output the bitstream. Simultaneously, the residual coefficients undergo inverse quantization and inverse transformation to obtain the reconstructed residual parameters. These reconstructed residual parameters are then added to the predicted values to obtain the reconstructed PUs. The reconstructed PUs form the F... n The corresponding reconstructed image, after being filtered within the loop, yields F. n Corresponding reconstructed frame F' n The reconstructed frame (i.e., the nth frame) enters the reference frame queue as a reference frame for subsequent video frames to be encoded, thus proceeding sequentially. The loop-in filtering process includes a deblocking filter (DB) algorithm and a sample adaptive bias algorithm (SAO). DB reduces blockiness, making the image more natural; SAO reduces distortion by adjusting pixel values. Intra-loop filtering is a step in the video encoding process, also including a deblocking filter (DB) algorithm and a sample adaptive bias algorithm (SAO). The deblocking filter reduces blockiness, making the image more natural; the sample adaptive bias algorithm reduces distortion by adjusting pixel values. During encoding, the reconstructed image, after intra-loop filtering, becomes a reconstructed frame, which enters the reference frame queue as a reference frame for subsequent video frames to be encoded.
[0087] In the process of dividing, F is first... nThe code is divided into at least one CTU (e.g., four). The size of the CTU can be 64×64, 32×32, or 16×16, and the specific size can be determined according to the encoding algorithm. Each CTU can be recursively divided into at least one CU using a quadtree. The size of the CU can be 64x64, 32x32, 16x16, or 8x8, and the size of the CU can only be less than or equal to the size of the CTU. Here, we take a 64x64 CTU as an example to introduce the process of dividing the CTU into CUs using a quadtree recursively: First, divide from top to bottom. As shown in Figure 5, starting from the level depth = 0, the 64x64 CTU is first divided into four 32x32 CUs. Then... Each 32x32 CU is divided into four 16x16 CUs, and so on, until depth=3, where the CU size is 8x8. Then, pruning is performed from bottom to top: the rate-distortion cost (RDcost) corresponding to the four 8x8 CUs (belonging to the same 16x16 CU) is calculated and summed to obtain cost1. The RDcost of the 16x16 CU to which the four 8x8 CUs belong is calculated to obtain cost2. If cost1 is less than cost2, the 8x8 CU segment is retained; otherwise, pruning continues upwards, comparing layer by layer, until a 64x64 CTU is obtained, thus obtaining at least one CU contained in the CTU. Rate-distortion cost is a metric used in video coding to measure the trade-off between coding efficiency and image quality. When dividing video frames (e.g., dividing CTU into CU, CU into PU, etc.) and selecting a prediction method, the rate-distortion cost corresponding to different division or prediction methods is calculated. The optimal division or prediction method is selected by comparing the rate-distortion costs, so as to reduce the bit rate required for encoding as much as possible while ensuring a certain image quality.
[0088] The prediction method for PU (Programming Process) can be intra-frame prediction or inter-frame prediction. Inter-frame prediction uses a reconstructed video frame as a reference frame, for example, the (n-1)th reconstructed frame, and performs prediction based on motion estimation (ME) and motion compensation (MC). Motion compensation is a technique that uses motion vectors obtained from motion estimation to extract prediction blocks from the reference frame to predict the pixel values to be encoded in the current frame. It is an important component of inter-frame prediction, utilizing the temporal correlation between video frames to eliminate temporal redundancy and achieve image compression.
[0089] When dividing a CU into PUs, different methods can be used to divide the CU. Then, the RDcost corresponding to each division method is calculated using multiple prediction methods. The division method with the smallest RDcost is selected to obtain the PUs contained in the CU. Then, the RDcost corresponding to each PU in multiple prediction methods is calculated, and the prediction method with the smallest RDcost is used as the final prediction method for the PU.
[0090] In the process of dividing CU into TU, CU can also be divided recursively using a quadtree. The process of dividing CU into TU using a quadtree recursively is similar to the process of dividing CTU into CU using a quadtree recursively, and will not be described in detail here.
[0091] Correspondingly, the types of enhanced video frame coding parameters include, but are not limited to, bitrate, quantization parameter (QP), prediction method (prediction type, prediction mode), coding method (inter-frame coding, intra-frame coding), and keyframe period (Group of Pictures, GOP). Here, GOP is the interval between two I-frames. An I-frame is an important frame type in video coding, also known as an intra-coded frame or keyframe, which uses intra-frame coding (a coding technique that utilizes intra-frame prediction and intra-frame compression). Enhanced video frame coding parameters can include frame-level coding parameters and unit-level coding parameters, such as the video frame's bitrate, the prediction method corresponding to the PU, the video frame's quantization parameter, and the quantization parameter corresponding to the coding unit. Specifically, the encoding parameters at the frame level corresponding to the enhanced video frame can be calculated based on the content enhancement information, or the encoding parameters of some units (e.g., subsequent optimization regions) corresponding to the enhanced video frame can be calculated based on the content enhancement information. The specific calculation method can be flexibly set according to actual needs. For example, the bitrate corresponding to the enhanced video frame can be positively correlated with the enhancement ratio corresponding to the enhanced video frame to ensure that the enhancement effect is preserved more after encoding. The enhancement ratio is the ratio between the area of the content-enhanced region in the enhanced video frame and the total area of the video frame.
[0092] In an optional implementation, the encoding parameters corresponding to the enhanced video frame can be calculated based on the bitrate control strategy corresponding to the video to be encoded. Based on content enhancement information, the encoding parameters corresponding to the enhanced video frame are optimized to obtain the enhanced video frame encoding parameters. The image quality corresponding to the optimized encoding parameters is higher than the image quality corresponding to the unoptimized encoding parameters. Specifically, the encoding parameters can be optimized at the frame level. For example, if the enhancement ratio of the enhanced video frame is high, the frame-level encoding parameters of the enhanced video frame can be optimized; if the enhancement ratio of the enhanced video frame is low, its encoding parameters do not need to be optimized. Alternatively, the encoding parameters of all or some units corresponding to the enhanced video frame can also be optimized.
[0093] In a specific example, the content enhancement information includes the area ratio A of the enhanced region, the average enhancement magnitude M, and the weighted value W of the region importance. The coding parameters of the enhanced video frame are calculated according to the following formulas: bitrate R = R0 × (1 + k1 × A + k2 × M + k3 × W), quantization parameter QP = QP0 - k4 × (A + M + W). Where R0 and QP0 are the initial bitrate and quantization parameter without considering content enhancement, respectively, and k1, k2, k3, and k4 are weighting coefficients determined experimentally.
[0094] Step S230: Encode the enhanced video frame according to the enhanced video frame coding parameters to obtain the encoded data of the enhanced video frame, and generate the encoded data of the video to be encoded based on the encoded data of the enhanced video frame.
[0095] Encoding enhanced video frames based on their encoding parameters yields their encoded data. Based on the encoded data of the enhanced video frames contained in the video to be encoded, the encoded data of the video to be encoded can be generated.
[0096] In the embodiment shown in Figure 2, on the one hand, content enhancement is performed on the video frame before encoding, which can improve the image quality of the encoded video frame; on the other hand, the encoding parameters of the video frame are calculated based on the content enhancement information, thereby deeply integrating the content enhancement process with the encoding process, reducing the weakening effect of the encoding process on the enhancement effect, improving the enhancement effect of the encoded video frame, reducing the loss of video frame details, further improving the quality of the encoded video frame, and improving the quality and visual effect of the encoded video.
[0097] In an exemplary embodiment, referring to FIG6, FIG6 is a flowchart of a video encoding method proposed based on FIG2. The method can be applied to the implementation environment shown in FIG1, and can be executed by the terminal device 110 in the implementation environment shown in FIG1, or by the server 120 in the implementation environment shown in FIG1, or can be executed jointly by the terminal device 110 and the server 120 in the implementation environment shown in FIG1.
[0098] As shown in Figure 6, step S220 may include steps S610-S630, which are described in detail below:
[0099] Step S610: Select the optimization region to be optimized for encoding parameters from the enhanced video frame, and generate content enhancement information based on the optimization region.
[0100] The content enhancement information includes information corresponding to the optimization region, such as the coordinates of the optimization region. The optimization region is selected from the enhanced video frame, and its encoding parameters need to be optimized. The specific selection method of the optimization region can be flexibly set according to actual needs. For example, considering that regions with large enhancement amplitude, regions with high video quality, regions of interest, and foreground regions are usually key areas of focus in video frames, in order to ensure the image quality such as clarity after encoding these regions, at least one of these regions can be selected from the enhanced video frame as the optimization region. Among them, the region with large enhancement amplitude can be a region whose enhancement amplitude parameter is greater than the enhancement amplitude threshold, and the region with high video quality can be a region whose video quality parameter is greater than the video quality threshold. For a detailed description of the enhancement amplitude parameter, enhancement amplitude threshold, video quality parameter, video quality threshold, and region of interest, please refer to the description in the subsequent embodiments, which will not be repeated here.
[0101] Step S620: Obtain the original coding parameters of the optimized region, and optimize the original coding parameters to obtain the optimized coding parameters of the optimized region; wherein, the image quality corresponding to the optimized coding parameters is higher than the image quality corresponding to the original coding parameters.
[0102] The original coding parameters refer to the coding parameters before optimization, while the optimized coding parameters are the coding parameters obtained by optimizing the original coding parameters. To ensure the image quality of the optimized region, the image quality corresponding to the optimized coding parameters is higher than that corresponding to the original coding parameters. Coding parameters can affect the quality of the encoded image, and the relationship between different types of coding parameters and image quality varies. Correspondingly, different optimization methods are used for different types of original coding parameters. For original coding parameters that are negatively correlated with image quality, the original coding parameters can be reduced to obtain optimized coding parameters; that is, the optimized coding parameters are smaller than the original coding parameters. For example, if the original coding parameters include original quantization parameters, since quantization parameters are negatively correlated with image quality, the smaller the quantization parameter, the finer the quantization and the higher the image quality. Conversely, the larger the quantization parameter, the more details are lost, the more image distortion is enhanced, and the quality decreases. Therefore, the original quantization parameters can be reduced to obtain optimized quantization parameters, which are smaller than the original quantization parameters. For the original coding parameters that are positively correlated with image quality, the original coding parameters can be increased to obtain optimized coding parameters, that is, the optimized coding parameters are greater than the original coding parameters. For example, if the original coding parameters include the original bitrate, since the bitrate is positively correlated with image quality, the original bitrate can be increased to obtain an optimized bitrate, which is greater than the original bitrate.
[0103] Step S630: Generate enhanced video frame coding parameters based on the optimized coding parameters of the optimized region.
[0104] The optimized coding parameters of the optimized region are added to the enhanced video frame coding parameters. In other words, the enhanced video frame coding parameters include the optimized coding parameters of the optimized region, so that the optimized region is encoded based on the optimized coding parameters to obtain the encoded data of the optimized region.
[0105] Optionally, the enhanced video frame encoding parameters may also include encoding parameters for other regions in the enhanced video frame besides the optimized region, so as to encode other regions based on the encoding parameters of other regions to obtain encoded data of other regions. The encoded data of the enhanced video frame includes the encoded data of the optimized region and other regions.
[0106] In one optional implementation, the original coding parameters of other regions can be added to the coding parameters of the enhanced video frame. That is, other regions are encoded based on their original coding parameters. In another optional implementation, after optimizing the original coding parameters of the optimized region, the bitrate and other resources occupied by the optimized region may be increased. In order to ensure the resource balance occupied by the enhanced video frame, the original coding parameters of other regions can be adjusted to obtain adjusted coding parameters. The adjusted coding parameters of other regions are then added to the coding parameters of the enhanced video frame. The image quality corresponding to the adjusted coding parameters is lower than that corresponding to the original coding parameters. The original coding parameters of all other regions can be adjusted, or only the original coding parameters of non-critical regions in other regions can be adjusted, such as edge regions and background regions.
[0107] It should be noted that the specific implementation details of steps S210 and S230 shown in Figure 6 can be found in steps S210 and S230 shown in Figure 2, and will not be repeated here.
[0108] In the embodiment shown in Figure 6, an optimization region is selected from the enhanced video frame, and the original coding parameters of the optimization region are optimized to obtain optimized coding parameters. Furthermore, the image quality corresponding to the optimized coding parameters is greater than the image quality corresponding to the original coding parameters. This allows more details of the optimization region to be preserved after encoding the optimization region based on the optimized coding parameters, thereby improving the image quality of the optimized region after encoding and reducing the weakening effect of the encoding process on the enhancement effect.
[0109] In an exemplary embodiment, referring to FIG7, FIG7 is a flowchart of a video encoding method proposed based on FIG6. The method can be applied to the implementation environment shown in FIG1, and can be executed by the terminal device 110 in the implementation environment shown in FIG1, or by the server 120 in the implementation environment shown in FIG1, or can be executed jointly by the terminal device 110 and the server 120 in the implementation environment shown in FIG1.
[0110] As shown in Figure 7, step S610 may include steps S710-S730, which are described in detail below:
[0111] Step S710: Calculate the enhancement amplitude parameter of each region contained in the enhanced video frame, and find the region in the enhanced video frame whose corresponding enhancement amplitude parameter is greater than the enhancement amplitude threshold.
[0112] The enhancement magnitude parameter is used to characterize the extent of enhancement during the content enhancement process, while the enhancement magnitude threshold is used to determine whether the enhancement magnitude is too large. The specific value can be flexibly set according to actual needs.
[0113] To prevent the loss of enhancement effects during encoding and compression, which could lead to a loss of detail in video frames, regions with larger enhancement magnitudes can be selected from the enhanced video frames as optimization regions. In the selection process, the enhanced video frames can be segmented to obtain at least two regions, and the enhancement magnitude parameter for each region can be calculated. Regions with enhancement magnitude parameters greater than the enhancement magnitude threshold can be selected as optimization regions.
[0114] The specific calculation method for the enhancement amplitude parameter and the specific selection method for the optimization region can be flexibly set according to actual needs, for example:
[0115] In an optional implementation, the change in image parameters of a region during the content enhancement process can reflect the enhancement magnitude of the region. Therefore, the change in image parameters of a region can be used as the enhancement magnitude parameter of the region. Correspondingly, the enhancement magnitude threshold includes the image parameter threshold. The process of calculating the enhancement magnitude parameter of each region contained in the enhanced video frame and finding the region in the enhanced video frame whose corresponding enhancement magnitude parameter is greater than the enhancement threshold may include:
[0116] According to the first size parameter, the original video frame and the enhanced video frame are segmented to obtain at least two original regions contained in the original video frame and at least two enhanced regions contained in the enhanced video frame; wherein, the enhanced regions correspond one-to-one with the original regions; based on the image parameters of each enhanced region and the image parameters of the original region corresponding to each enhanced region, the image parameter change value corresponding to each enhanced region is calculated, and the image parameter change value is used as the enhancement amplitude parameter of each enhanced region; from the at least two enhanced regions, the enhanced region whose corresponding image parameter change value is greater than the image parameter threshold is found.
[0117] The first size parameter is the size parameter corresponding to a single original region or a single enhanced region. Its specific value can be flexibly set according to actual needs. For example, the minimum size of any type of unit among CU, PU, and TU can be used as the first size parameter. For example, the minimum size of TU, 4×4, can be used as the first size parameter.
[0118] The original region is a region in the original video frame, and the enhanced region is a region in the enhanced video frame. The original and enhanced video frames are segmented in the same way; therefore, there is a one-to-one correspondence between the original and enhanced regions. The enhanced region corresponding to the original region is the region obtained after content enhancement of that original region. Image parameters describe pixel parameters such as color and brightness, and their format can be YUV data, RGB data, etc. Comparing the image parameters of the original region with the image parameters of the corresponding enhanced region (equivalent to comparing the image parameters of the same region in the original video frame with those in the enhanced video frame) yields the change value of the image parameters for that enhanced region. This change value is used as the enhancement amplitude parameter for that enhanced region. The absolute value of the difference between the image parameters of the original region and the corresponding enhanced region can be used as the image parameter change value. Furthermore, the selected enhanced region's image parameter change value must be greater than the image parameter threshold.
[0119] In another optional implementation, the video quality of the enhanced region can also reflect the enhancement magnitude of the region. Therefore, the enhancement magnitude parameter of the enhanced region can be determined based on the video quality of the enhanced region. Correspondingly, the enhancement magnitude threshold includes a video quality threshold. The process of calculating the enhancement magnitude parameter of each region contained in the enhanced video frame and finding the region in the enhanced video frame whose corresponding enhancement magnitude parameter is greater than the enhancement magnitude threshold may include:
[0120] According to the second size parameter, the enhanced video frame is segmented to obtain at least two candidate regions contained in the enhanced video frame; video quality detection is performed on each candidate region to obtain the video quality parameter of each candidate region, and the enhancement amplitude parameter of each candidate region is calculated based on the video quality parameter of each candidate region; and candidate regions with corresponding video quality parameters greater than the video quality threshold are found from the at least two candidate regions.
[0121] The second size parameter is the size parameter of a single candidate region. The second size parameter can be the same as or different from the first size parameter. The second size parameter can be larger than the first size parameter. The specific value can be flexibly set according to actual needs. For example, the maximum size of the CU can be used as the second size parameter. For example, in the video coding standard H.266, the maximum size of the CU is 128×128, so 128×128 can be used as the second size parameter.
[0122] Video quality inspection is used to perceive, measure, and evaluate the content and image quality of video frames to obtain video quality parameters that characterize video quality. Specific detection methods include, but are not limited to, video quality assessment (VQA).
[0123] Video quality detection is used to perceive, measure, and evaluate the content and image quality of video frames, obtaining video quality parameters that characterize video quality. Specific detection methods include, but are not limited to, Video Quality Assessment (VQA). Taking the VQA method based on the Structural Similarity Index (SSIM) as an example, the SSIM is an indicator used to measure the degree of structural and content similarity between two image or video regions, and is commonly used in video quality assessment. By calculating the mean, variance, covariance, and other parameters of the two regions, and introducing constants to avoid zero denominators, the final SSIM value is obtained. The closer the value is to 1, the more similar the structure and content of the two regions are, and the higher the video quality. When performing video quality detection on candidate regions, for each candidate region, it is compared with the corresponding region in the original video frame. The specific calculation process is as follows: Calculate the mean μ of the two regions. x and μ y , Where x and y represent the pixel matrices of the original region and the candidate region, respectively, and M and N are the width and height of the region. The variance between the two regions is calculated. and Calculate the covariance σ of the two regions xy , Two constants, C1 and C2, are introduced to avoid the case where the denominator is zero. Usually, C1 = (K1 × L). 2 C2 = (K2 × L) 2 K1 = 0.01, K2 = 0.03, and L is the dynamic range of the pixel value (e.g., 255). Calculate the SSIM value. The closer the SSIM value is to 1, the more similar the structure and content of the candidate region is to the original region, and the higher the video quality. This SSIM value is used as a video quality parameter for the candidate region.
[0124] The enhancement magnitude parameter of a candidate region is calculated based on the video quality parameters of that region. The specific calculation method can be flexibly set according to actual needs. In one optional example, the video quality parameters of the candidate region can be directly used as the enhancement magnitude parameter. In another optional example, the region corresponding to the candidate region can be found in the original video frame, and video quality detection can be performed on that region to obtain the pre-enhancement video quality parameters of the candidate region. The difference (or the absolute value of the difference) between the video quality parameters of the candidate region and the pre-enhancement video quality parameters can be used as the enhancement magnitude parameter of the candidate region. In yet another optional example, the process of calculating the enhancement magnitude parameter of each candidate region based on its video quality parameters may include: performing video quality detection on the enhanced video frame to obtain the video quality parameters of the enhanced video frame; and performing a difference operation between the video quality parameters of each candidate region and the video quality parameters of the enhanced video frame to obtain the enhancement magnitude parameter of each candidate region.
[0125] The difference operation can be a difference operation or an absolute difference operation. That is to say, the enhancement amplitude parameter of the candidate region is the difference between the video quality parameter of the candidate region and the video quality parameter of the enhanced video frame, or the absolute value of the difference.
[0126] Step S720: The found region is used as the optimization region.
[0127] The optimization region includes the areas in the enhanced video frames where the corresponding enhancement amplitude parameter is greater than the enhancement amplitude threshold.
[0128] Step S730: Generate content enhancement information based on the optimized region.
[0129] The specific process for generating content-enhanced information can be found in the description of the foregoing embodiments, and will not be repeated here.
[0130] It should be noted that the specific implementation details of steps S210 and S230 shown in Figure 7 can be found in steps S210 and S230 shown in Figure 2, and the specific implementation details of steps S620-S630 shown in Figure 7 can be found in steps S620-S630 shown in Figure 6. They will not be repeated here.
[0131] In the embodiment shown in Figure 7, the region with a higher enhancement magnitude is taken as the optimization region, and its corresponding original encoding parameters are optimized to improve the image quality of the optimized region after encoding. This reduces the weakening effect of the encoding process on the enhancement effect, improves the enhancement effect of the encoded video frame, and further improves the image quality of the encoded video frame.
[0132] In an exemplary embodiment, referring to FIG8, FIG8 is a flowchart of a video encoding method proposed based on FIG7. The method can be applied to the implementation environment shown in FIG1, and can be executed by the terminal device 110 in the implementation environment shown in FIG1, or by the server 120 in the implementation environment shown in FIG1, or can be executed jointly by the terminal device 110 and the server 120 in the implementation environment shown in FIG1.
[0133] As shown in Figure 8, step S710 may include steps S810-S830, which are described in detail below:
[0134] Step S810: Calculate the enhancement amplitude parameters of each region contained in the enhanced video frame, and obtain the video quality requirement information corresponding to the video to be encoded.
[0135] The specific calculation process for the enhancement amplitude parameter can be flexibly set according to actual needs. Video quality requirement information is used to characterize the desired video quality, including but not limited to resolution requirements, such as high-definition video and low-definition video.
[0136] Step S820: Calculate the enhancement amplitude threshold corresponding to the video to be encoded based on the video quality requirement information; wherein, the enhancement amplitude threshold is negatively correlated with the video quality requirement information.
[0137] Different video quality requirements dictate different video frame quality standards. Therefore, an enhancement threshold can be set that is negatively correlated with the video quality requirements of the video to be encoded. In other words, the higher the video quality requirement, the lower the enhancement threshold, resulting in more regions being selected for optimization and improving the video quality of the encoded video frame. Conversely, the lower the video quality requirement, the higher the enhancement threshold, resulting in fewer regions being selected for optimization, reducing the computational resources used in the video frame encoding process and improving video encoding efficiency.
[0138] Optionally, a mapping table containing enhancement amplitude thresholds corresponding to different video quality levels can be pre-set, so that the enhancement amplitude threshold corresponding to the video to be encoded can be found from the mapping table according to the video quality level to which the video quality requirement information of the video to be encoded belongs; or, a negative correlation mapping algorithm between video quality requirement information and enhancement amplitude threshold can be pre-set, so that the video quality requirement information is used as input and the enhancement amplitude threshold corresponding to the video to be encoded is calculated by the algorithm.
[0139] For example, the video quality threshold corresponding to high-definition video can be greater than or equal to 10, the video quality threshold corresponding to low-definition video can be greater than or equal to 5, the image parameter threshold corresponding to high-definition video can be greater than or equal to 30%, and the image parameter threshold corresponding to low-definition video can be greater than or equal to 10%. Here, high-definition video can be a video with a resolution greater than a first resolution threshold, and low-definition video can be a video with a resolution less than a second resolution threshold, where the first resolution threshold is greater than or equal to the second resolution threshold.
[0140] Step S830: Find the region in the enhanced video frame whose corresponding enhancement amplitude parameter is greater than the enhancement amplitude threshold corresponding to the video to be encoded.
[0141] The enhancement magnitude parameter of the found region is greater than the enhancement magnitude threshold corresponding to the video to be encoded to which the region belongs.
[0142] It should be noted that the specific implementation details of steps S210 and S230 shown in Figure 8 can be found in steps S210 and S230 shown in Figure 2, the specific implementation details of steps S620-S630 shown in Figure 8 can be found in steps S620-S630 shown in Figure 6, and the specific implementation details of steps S720-S730 shown in Figure 8 can be found in steps S720-S730 shown in Figure 7. They will not be repeated here.
[0143] In the embodiment shown in Figure 8, different enhancement thresholds correspond to different video quality requirements, and the enhancement thresholds are negatively correlated with the video quality requirements. This results in more regions in video frames with high video quality requirements being selected as optimization regions, and fewer regions in video frames with low video quality requirements being selected as optimization regions. This saves computational resources during the encoding process and improves video encoding efficiency while meeting video quality requirements.
[0144] In an exemplary embodiment, referring to FIG9A, FIG9A is a flowchart of a video encoding method proposed based on FIG6. The method can be applied to the implementation environment shown in FIG1, and can be executed by the terminal device 110 in the implementation environment shown in FIG1, or by the server 120 in the implementation environment shown in FIG1, or can be executed jointly by the terminal device 110 and the server 120 in the implementation environment shown in FIG1.
[0145] As shown in Figure 9A, step S620 may include steps S910-S930, which are described in detail below:
[0146] Step S910: Obtain the original encoding parameters of the optimized region.
[0147] The specific method for obtaining the original encoding parameters can be found in the description of the subsequent embodiments, and will not be repeated here.
[0148] Step S920: From the parameter optimization methods corresponding to various region types, find the target parameter optimization method corresponding to the optimized region; where the parameter optimization methods corresponding to different region types are not completely the same.
[0149] Parameter optimization methods are used to optimize the original encoding parameters. In an optional implementation, considering that the optimization regions have different region types and attributes, the corresponding parameter optimization methods may not be completely identical. Different parameter optimization methods may have at least one different optimization content; for example, the type of encoding parameter being optimized, the optimization magnitude, etc. For instance, parameter optimization method 1 optimizes both the bitrate and quantization parameters; parameter optimization method 2 optimizes only the quantization parameters; parameter optimization method 3 optimizes the bitrate by 20%; and parameter optimization method 4 optimizes the bitrate by 30%, etc. The target parameter optimization method refers to the parameter optimization method corresponding to the region type to which the optimization region belongs. It should be noted that in other optional implementations of this application, in order to simplify the parameter optimization process, improve parameter optimization efficiency, and save resources occupied during parameter optimization, optimization regions of different region types can adopt the same parameter optimization method.
[0150] The parameter optimization method for each region type can be set according to the importance, attributes, and application scenarios of that region type. The specific classification method and classification level can be set according to the application scenario. The optimization magnitude can be positively correlated with the importance of the region type. For example, for enhancement region AREA1, where the image parameter change value is greater than the image parameter threshold, the image parameter change is large, and to preserve the enhancement effect, its corresponding optimization magnitude can be large. For candidate region AREA2, where the video quality parameters are greater than the video quality threshold, the bitrate occupied by its original encoding parameters is large; if its optimization magnitude is too large, it may lead to uncontrollable video output bitrate, therefore, its optimization magnitude can be small. For region of interest AREA3, which is usually a relatively important region, the optimization magnitude can be large. For example, referring to Figure 9B, after content enhancement, the image parameter change of the region to which the "tree" belongs is large, its type is AREA1, and the optimization magnitude is 10%; the video quality of the region to which the "person" belongs is high, its type is AREA2, and the corresponding optimization magnitude is 2%; the region shown by the "car" is a region of interest, its type is AREA3, and the optimization magnitude is 12%.
[0151] Optionally, a Region of Interest (ROI) refers to a region that receives special attention or processing in image processing. If the ROI is a region obtained by object detection in an enhanced video frame, then different types of target objects belong to the ROI, and therefore, different target object types have different importance. Thus, the ROI can be further classified based on the importance of the target object, and various parameter optimization methods can be set for each ROI. The optimization magnitude can be positively correlated with the importance of the target object. Target objects include, but are not limited to, people, animals, game heroes, moving objects, and text. The total number of categories, classification levels, and the importance of the target object corresponding to the ROI can be set according to the specific application scenario. In an optional example, if the video is a game video, the regions belonging to faces, hair, and game heroes can be classified as the first category, the regions belonging to moving objects and text as the second category, and other ROIs as other categories. Furthermore, since the importance of the first type of ROI > the importance of the second type of ROI > the importance of the third type of ROI, the optimization magnitude of the first type of ROI > the optimization magnitude of the second type of ROI > the optimization magnitude of the third type of ROI.
[0152] Step S930: Optimize the original coding parameters of the optimization region according to the target parameter optimization method to obtain the optimized coding parameters of the optimization region.
[0153] Based on the target parameter optimization method corresponding to the optimization region, the original coding parameters of the optimization region are optimized to obtain the optimized coding parameters.
[0154] In a specific example, depending on the target parameter optimization method, if the optimized region is a region of interest obtained by object detection of the enhanced video frame, the importance of its region type is I. For the original bitrate R... region Optimize bitrate R optimized =R region ×(1+k8×I); for the original quantization parameter QP region Optimize quantization parameter QP optimized =QP region -k9×I, where k8 and k9 are the bitrate and quantization parameter adjustment coefficients related to importance, respectively.
[0155] It should be noted that the specific implementation details of steps S210 and S230 shown in Figure 9A can be found in steps S210 and S230 shown in Figure 2, and the specific implementation details of steps S610 and S630 shown in Figure 9A can be found in steps S610 and S630 shown in Figure 6. They will not be repeated here.
[0156] In the embodiment shown in Figure 9A, different optimization regions correspond to different parameter optimization methods, thereby meeting different business needs and improving the flexibility of parameter optimization.
[0157] In an exemplary embodiment, referring to FIG10, FIG10 is a flowchart of a video encoding method proposed based on FIG9A. The method can be applied to the implementation environment shown in FIG1, and can be executed by the terminal device 110 in the implementation environment shown in FIG1, or by the server 120 in the implementation environment shown in FIG1, or can be jointly executed by the terminal device 110 and the server 120 in the implementation environment shown in FIG1.
[0158] As shown in Figure 10, step S930 may include steps S1010-S1030, which are described in detail below:
[0159] Step S1010: Obtain the original quantization parameters of the optimized region from the original encoding parameters of the optimized region.
[0160] The type of encoding parameter that needs to be optimized is a quantization parameter. Correspondingly, the original encoding parameters of the optimization region include quantization parameters, that is, the original quantization parameters.
[0161] Step S1020: Based on the target parameter optimization method, reduce the original quantization parameter of the optimization region to obtain the optimized quantization parameter of the optimization region; wherein, the reduction of quantization parameter corresponding to different parameter optimization methods is not exactly the same.
[0162] The quantization parameter directly affects image quality, and the impact of other types of coding parameters on image quality is based on their influence on the quantization parameter, thus being implemented through the quantization parameter. Therefore, during parameter optimization, the quantization parameter can be directly optimized, avoiding the need to optimize other types of coding parameters and then calculate the quantization parameter based on the optimized coding parameters. Since the quantization parameter is negatively correlated with image quality, to ensure that the image quality corresponding to the optimized quantization parameter is greater than that of the original quantization parameter, the original quantization parameter can be reduced to obtain the optimized quantization parameter. In other words, the optimized quantization parameter is less than or equal to the original quantization parameter.
[0163] Specifically, the reduction rate of the original quantization parameters can vary for different types of optimized regions. This reduction rate can be determined based on the attributes and importance of the optimized region itself. In an optional example, the process of reducing the original quantization parameters of the optimized region to obtain its optimized quantization parameters, according to the target parameter optimization method, can include: if the optimized region is a region of interest obtained by target detection of the enhanced video frame, then the target object to which the optimized region belongs is obtained; the original quantization parameters of the optimized region are reduced according to the importance of the target object to obtain its optimized quantization parameters; wherein the reduction rate of the original quantization parameters of the optimized region is positively correlated with the importance of the target object.
[0164] In other words, the more important the target object, the greater the reduction in its original quantization parameters. This reduces loss in the region containing the target object during encoding, improving the image quality of video frames. Conversely, the less important the target object, the smaller the reduction in its original quantization parameters. This reduces the computational resource consumption during video encoding, improving efficiency. The reduction magnitude can be characterized by numerical or proportional reductions. For example, the reduction magnitude for the first type of ROI can be set to at least 3%; for the second type of ROI, at least 2%; and for other types of ROI, at least 1%. It should be noted that the reduction magnitudes for each region are merely illustrative; the specific values can be set based on the actual application scenario and the allowable video bitrate fluctuation range.
[0165] Step S1030: Based on the optimized quantization parameters of the optimized region, generate optimized coding parameters for the optimized region.
[0166] Add the optimized quantization parameters to the optimized encoding parameters of the optimized region.
[0167] It should be noted that the specific implementation details of steps S210 and S230 shown in Figure 10 can be referred to steps S210 and S230 shown in Figure 2, the specific implementation details of steps S610 and S630 shown in Figure 10 can be referred to steps S610 and S630 shown in Figure 6, and the specific implementation details of steps S910-S920 shown in Figure 10 can be referred to steps S910-S920 shown in Figure 9A, and will not be repeated here.
[0168] In the embodiment shown in Figure 10, the original coding parameters are optimized by directly reducing the original quantization parameters according to the target parameter optimization method. This can reduce the complexity of parameter optimization, save computing resources, and improve video coding efficiency.
[0169] In an exemplary embodiment, referring to FIG11, FIG11 is a flowchart of a video encoding method proposed based on FIG6. The method can be applied to the implementation environment shown in FIG1, and can be executed by the terminal device 110 in the implementation environment shown in FIG1, or by the server 120 in the implementation environment shown in FIG1, or can be executed jointly by the terminal device 110 and the server 120 in the implementation environment shown in FIG1.
[0170] As shown in Figure 11, step S620 may include steps S1110-S1130, which are described in detail below:
[0171] Step S1110: Obtain the bitrate control strategy corresponding to the video to be encoded, and calculate the encoding parameters corresponding to the enhanced video frame according to the bitrate control strategy.
[0172] To prevent the actual bitrate of a video from exceeding the target bandwidth, which would prevent the encoded video data from being transmitted and thus render it unusable in a real-world environment, a corresponding bitrate control strategy can be set to control the bitrate during the video encoding process, ensuring that the actual bitrate of the video is lower than the target bandwidth. This bitrate control strategy includes, but is not limited to, at least one of the following:
[0173] Constant Quantization Parameter (CQP): This aims to maintain constant quantization distortion, meaning that each video frame is encoded using the same quantization parameter.
[0174] Constant Rate Factor (CRF): This controls the trade-off between quality and bit rate during the encoding process by setting a constant quality factor (i.e., CRF value).
[0175] Specify Average Bitrate (ABR): By setting a target average bitrate, the average bitrate of the entire video is made as close as possible to the target average bitrate during the encoding process;
[0176] Constant Bitrate (CBR): Requires the encoder to maintain a constant bitrate during the encoding process, meaning that the amount of data per unit time remains constant.
[0177] Video Buffering Verifier (VBV): An important mechanism used in video encoding and decoding to simulate and verify the behavior of the video stream in the decoder buffer. It helps the encoder generate a video stream that meets specific bitrate requirements and ensures that the buffer does not overflow or underflow, thus guaranteeing smooth video playback.
[0178] Two-Pass Encoding: A video coding strategy in which the video content is first encoded once to collect statistical information, and then this information is used to optimize coding parameters, such as quantization parameters, during the second encoding.
[0179] Based on the bitrate control strategy corresponding to the video to be encoded, the encoding parameters corresponding to the enhanced video frames can be calculated. The specific calculation method for these encoding parameters can be flexibly set according to actual needs.
[0180] In an optional implementation, the spatiotemporal complexity corresponding to the enhanced video frame can be calculated, and bitrate control can be performed based on the spatiotemporal complexity of the enhanced video frame to calculate the quantization parameters of the enhanced video frame. Based on the quantization parameters of the enhanced video frame, the quantization parameters of the coding units are then determined. For example, the bitrate control process can be seen in Figure 12A. For the input video requiring bitrate control, each frame of the input video is pre-analyzed to obtain the spatiotemporal complexity of the video frame. Bitrate control processing is performed based on the spatiotemporal complexity to obtain the quantization parameters (QP) of the video frame. Then, based on the quantization parameters of the video frame, the quantization parameters of each unit in the video frame are determined. The actual bitrate of the video frame is then calculated based on the quantization parameters of each unit, and the actual bitrate of the video frame and the target bitrate are cached.
[0181] Spatiotemporal complexity measures the spatial and temporal complexity of a video frame. It is a crucial factor in bitrate control during video coding. For example, when calculating the coding parameters for an enhanced video frame, bitrate control is performed based on the frame's spatiotemporal complexity to determine quantization parameters and other coding parameters. Spatiotemporal complexity can be calculated by transforming and summing the residual parameters of video frames. For instance, calculating the spatiotemporal complexity of a P-frame involves calculating the weighted total spatiotemporal complexity across multiple frames and the number of frames.
[0182] In a specific example, based on the bitrate control strategy corresponding to the video to be encoded, if the Constant Quantization Parameter (CQP) strategy is adopted, the quantization parameter QP corresponding to the enhanced video frame is directly set to a fixed value QP. fixed If a constant quality factor (CRF) strategy is adopted, the quantization parameter QP = CRF - k5 × C is calculated based on the spatiotemporal complexity C of the enhanced video frame, where k5 is a coefficient related to the quality factor and complexity. If a specified average bitrate (ABR) strategy is adopted, the initial bitrate R is first calculated based on the spatiotemporal complexity C of the enhanced video frame. init =R target ×C / C avg Then adjust the quantization parameter QP according to the initial bitrate, QP = QP base -k6×log2(R init / R base ), where Rtarget It is the target average bitrate, C avg It is the average time and space complexity of all video frames, QP base and R base These are the base quantization parameters and base bitrate; k6 is the adjustment coefficient.
[0183] In another optional implementation, to improve the accuracy of bitrate control, the initial quantization parameters of the enhanced video frame can be calculated based on the frame type and spatiotemporal complexity. The initial quantization parameters are then adjusted based on the actual total number of bits consumed by the encoded video frames preceding the enhanced video frame in the video to be encoded, as well as the pre-allocated total number of bits. The quantization parameters of the enhanced video frame are then calculated based on the adjusted initial quantization parameters. The frame types include I-frames, B-frames, and P-frames. B-frames are bidirectional difference frames, and P-frames are forward predictive coded frames. For example, the quantization parameters of the enhanced video frame can be calculated using the method shown in Figure 12B. As shown in Figure 12B, the calculation process of the video frame's quantization parameters mainly includes the following sections 1.1-1.3, which are detailed below:
[0184] 1.1 Calculate the initial quantization parameters for the current frame.
[0185] The video encoding process involves encoding video frames sequentially. When encoding the current frame, the frame type is first determined, and then the initial quantization parameters for the current video frame are calculated using different methods based on the frame type.
[0186] Calculation of initial quantization parameters for an I-frame: The initial quantization parameters of an I-frame can be the average quantization parameters of the encoded video frames. That is, the initial quantization parameters of an I-frame = the sum of the quantization parameters of the encoded video frames ÷ the number of encoded video frames. For the first I-frame, since there are no encoded video frames, its initial quantization parameters can be a preset value (e.g., 24).
[0187] Optionally, in calculating the sum of quantization parameters for encoded video frames, if the encoded video frames contain video frames of different frame types, it is necessary to map the quantization parameters of the encoded video frames to quantization parameters of the same frame type according to ipratio and pbratio before performing the summation operation. Here, ipratio is the increment of the quantizer for I-frames compared to the quantizer for P-frames, resulting in higher image quality for I-frames; its specific value can be 1.4, etc. Pbratio is the decrement of the quantizer for B-frames compared to the quantizer for P-frames, resulting in lower image quality for B-frames. The mapping method for quantization parameters of I-frames to quantization parameters of P-frames is: Quantization parameters of P-frames = Quantization parameters of I-frames + 6.0. *log2(ipratio); The way to map the quantization parameters of a B-frame to the quantization parameters of a P-frame is: quantization parameter of P-frame = quantization parameter of B-frame - 6.0 * log2(pbratio); For example, if the encoded video frames include frames 1-3, frames 2 and 3 are P-frames, and frame 1 is not a P-frame, then the quantization parameters of frame 1 can be mapped to the quantization parameters of P-frames, and then added to the quantization parameters of frames 2 and 3. Wherein, if frame 1 is an I-frame, then the mapped quantization parameter = quantization parameter of frame 1 + 6.0 * log2(ipratio); if frame 1 is a B-frame, then the mapped quantization parameter = quantization parameter of frame 1 - 6.0 * log2(pbratio).
[0188] Calculation of initial quantization parameters for P-frames: First, calculate the time and space complexity of the P-frame, which is calculated as follows: cplxsum[i+1]=cplxsum[i]*0.5+SATD[i] cplxcount[i+1]=cplxcount[i]*0.5+1
[0189] Where cplxsum[i+1] refers to the weighted total time and space complexity from frame 1 to frame i+1, cplxsum[i] refers to the weighted total time and space complexity from frame 1 to frame i, and the initial value of cplxsum, i.e., cplxsum[0], can be 0; SATD[i] refers to the SATD (Sum of Absolute Transformed Difference) of frame i, that is, the residual parameters corresponding to frame i are transformed by Hadamard and then summed in absolute value. Here, frame i can be downsampled in lookahead and its cost can be calculated (the calculation method can be to perform simple intra-frame and inter-frame prediction on frame i, calculate the SATD of frame i under intra-frame prediction based on the intra-frame prediction result, calculate the SATD of frame i under inter-frame prediction based on the inter-frame prediction result, and select the optimal SATD from the SATD corresponding to intra-frame prediction and inter-frame prediction respectively) to obtain the SATD of frame i. blurred_complexity[i] refers to the spatiotemporal complexity of the i-th video frame, which is the weighted average spatiotemporal complexity from frame 1 to frame i. Hadamard transform is a transform coding method in video coding, and Lookahead is a key module in video coding used to analyze and predict future video frames to optimize the coding decision of the current frame.
[0190] Then, the initial quantization parameters of the P-frame are calculated as follows: qscale_raw[i] = blurred_complexity[i] / 1 - qcompress
[0191] Where qcompress is the compression factor, and the specific value can be set according to actual needs, for example, it can be 0.6; wanted_bits_window[i] refers to the cumulative value of the target number of bits from frame 0 to frame i, that is, the total target number of bits; qscale_adjust[i] is the initial quantization parameter of frame i.
[0192] Initial quantization parameter calculation for B-frame: Locate the reference frame for the B-frame. Select a reference value from ipratio and pbratio based on the frame type of the reference frame. Calculate the adjusted quantization parameter of the reference frame based on this reference value. Adjust the quantization parameter of the reference frame based on the adjusted quantization parameter to obtain the initial quantization parameter of the B-frame. Here, the adjusted quantization parameter = 6.0 * log2(x), where x is ipratio or pbratio. If there are two reference frames, calculate the initial quantization parameter using both reference frames separately. Average the two initial quantization parameters to obtain the initial quantization parameter of the B-frame.
[0193] Step 1.2, Adjustment of initial quantization parameters.
[0194] The adjustment formula is as follows: abr_buffer[i]=2*R T *sqrt(T total ) qscale_adjust[i]′=qscale_adjust[i]*overflow[i]
[0195] Among them, R T For the target bitrate, sqrt(T) total The total encoding time is from frame 1 to frame (i-1). The clip3 function is used to limit the output within the corresponding range. total_bits[i-1] is the actual total number of bits consumed from frame 1 to frame (i-1), wanted_bits[i-1] is the pre-allocated total number of bits from frame 1 to frame (i-1), and qscale_adjust[i]′ is the adjusted quantization parameter for frame i. This allows for adjustment of the initial quantization parameters through overflow adjustment, ensuring that the actual number of bits consumed and the pre-allocated number of bits are close throughout the encoding process. Optionally, overflow adjustment can be performed on the initial quantization parameters of P-frames, B-frames, and I-frames; or, only the initial quantization parameters of P-frames and B-frames can be overflow adjusted. Before overflow adjustment of the initial quantization parameters of B-frames, it can be determined whether the corresponding rate control strategy is CBR. If it is, overflow adjustment is performed; otherwise, overflow adjustment is not performed. Optionally, after performing Overflow adjustment on the P-frame, the quantization parameters of adjacent frames can be limited, that is, the maximum change in quantization parameters between adjacent frames can be set to prevent the quantization parameters from suddenly increasing or decreasing.
[0196] Step 1.3, Quantization parameter calculation.
[0197] The quantization parameters are calculated using the following formula:
[0198] Where QP[i] is the quantization parameter of the i-th frame.
[0199] Step S1120: Calculate the original coding parameters of the optimized region based on the coding parameters corresponding to the enhanced video frame.
[0200] Based on the coding parameters corresponding to the enhanced video frame, the coding parameters of the optimized region can be calculated; these coding parameters are the original coding parameters. In a specific example, based on the coding parameters corresponding to the enhanced video frame, if the coding parameters corresponding to the enhanced video frame include the bitrate R... frame and quantization parameter QP frameFor the optimized region, its original bitrate R region =R frame ×S region / S frame S region S is the area of the optimization region. frame It is the total area of the video frames; its original quantization parameter QP region =QP frame +k7×(1-S region / S frame k7 is an adjustment coefficient related to the area ratio of the region.
[0201] In an optional implementation, if the encoding parameters of the enhanced video frame include the quantization parameters of the enhanced video frame, then the quantization parameters of the enhanced video frame can be directly used as the original quantization parameters for the optimized region and other regions in the frame.
[0202] Step S1130: Optimize the original coding parameters to obtain the optimized coding parameters for the optimized region.
[0203] The specific optimization process can be found in the description of the foregoing embodiments, and will not be repeated here.
[0204] It should be noted that the specific implementation details of steps S210 and S230 shown in Figure 11 can be found in steps S210 and S230 shown in Figure 2, and the specific implementation details of steps S610 and S630 shown in Figure 11 can be found in steps S610 and S630 shown in Figure 6. They will not be repeated here.
[0205] In the embodiment shown in Figure 11, the original encoding parameters of the optimization region are calculated based on the bitrate control strategy corresponding to the video to be encoded. Optimizing these original encoding parameters can deeply integrate the content enhancement process with the bitrate control process, thereby ensuring the bitrate control effect while reducing the impact of bitrate control on the enhancement effect and improving the enhancement effect and image quality of the encoded video.
[0206] In an exemplary embodiment, referring to FIG13, FIG13 is a flowchart of a video encoding method proposed based on FIG6. The method can be applied to the implementation environment shown in FIG1, and can be executed by the terminal device 110 in the implementation environment shown in FIG1, or by the server 120 in the implementation environment shown in FIG1, or can be executed jointly by the terminal device 110 and the server 120 in the implementation environment shown in FIG1.
[0207] As shown in Figure 13, if the number of optimization regions includes at least two, the video coding method may further include steps S1310-S1320 before step S620, as detailed below:
[0208] Step S1310: Find the optimization regions that have overlapping regions from at least two optimization regions.
[0209] If at least two optimization regions are selected from the enhanced video frames, it can be determined whether there is any overlap between these optimization regions, so as to find the optimization regions with overlapping regions.
[0210] Step S1320: Merge the optimization regions with overlapping areas to obtain the merged optimization region.
[0211] For overlapping optimization regions, they are merged to obtain merged optimization regions. Then, based on the merged optimization regions and other optimization regions, subsequent processing is performed. For example, suppose optimization regions 1-5 are detected from any original video frame. Optimization regions 1 and 2 overlap, optimization regions 2 and 3 overlap, optimization region 4 does not overlap with optimization regions 1-3 and 5, and optimization region 5 does not overlap with optimization regions 1-4. Therefore, optimization regions 1-3 are merged to obtain merged optimization region 11. In subsequent processing, the original encoding parameters corresponding to optimization regions 11, 4, and 5 are obtained and optimized.
[0212] Optionally, considering that the parameter optimization methods for optimization regions of different region types may be different, only optimization regions with overlapping regions and belonging to the same region type can be merged. Continuing with the previous example, if optimization region 1 and optimization region 2 have the same region type, and optimization region 2 and optimization region 3 have different region types, then only optimization region 1 and optimization region 2 are merged to obtain the merged optimization region 12. In subsequent processing, the original encoding parameters corresponding to optimization regions 12, 3, 4, and 5 are obtained and the original encoding parameters are optimized.
[0213] It should be noted that the specific implementation details of steps S210 and S230 shown in Figure 13 can be referred to steps S210 and S230 shown in Figure 2, and the specific implementation details of steps S610 and S630 shown in Figure 13 can be referred to steps S610 and S630 shown in Figure 6. They will not be repeated here.
[0214] In the embodiment shown in Figure 13, when the number of optimization regions includes at least two, merging overlapping optimization regions and then optimizing parameters based on the optimization regions can save computing resources and improve video coding efficiency.
[0215] To better understand the present invention, a specific example is provided here. Referring to Figure 14, which is a schematic diagram of the video encoding process provided in this embodiment, the video encoding process includes:
[0216] Acquire video sources, supporting various standard containers and encoding formats as input.
[0217] Image quality enhancement: Decode the video source to obtain YUV data 1 of the video frame, and enhance the image quality of YUV data 1 to obtain YUV data 2. Image quality enhancement processing methods include traditional CV models, deep learning models, and the latest LLM+DiT large model, etc. For specific enhancement operations, please refer to the description in the foregoing embodiments, which will not be repeated here.
[0218] Then, based on YUV data 1 and YUV data 2, the areas with significant enhancement and restoration were analyzed. For example, in Figure 14, the area corresponding to the dashed box in the enhanced image was analyzed. The specific search process included:
[0219] 1. Locate region AREA1 and analyze and compare YUV data 1 and YUV data 2. The analysis and comparison are performed in 4×4 units. Compare YUV data 1 and YUV data 2 to find the region where the absolute value of the difference between the YUV data is greater than the threshold A, and denote it as AREA1. Different thresholds A can be set for videos of different quality. The higher the video quality, the lower the corresponding threshold A. For example, for high-definition videos, the threshold A can be 10%, and for low-definition videos, the threshold A can be 30%. The specific value of the threshold A can be flexibly configured according to the actual application scenario.
[0220] 2. Locate AREA2. Based on the YUV data 2 of the video frame, perform VQA detection to obtain the VQA parameters of the video frame. Then, perform regional VQA detection on the YUV data 2 of the video frame in 128×128 units to obtain the VQA parameters of each region. Calculate the absolute value of the difference between the VQA parameters of each region and the VQA parameters of the video frame. Regions with an absolute value of the difference greater than threshold B are identified as AREA2. Different thresholds B can be set for videos of different quality levels. The higher the video quality, the lower the corresponding threshold B. For example, for high-definition videos, the threshold B can be 5 points, and for low-definition videos, the threshold B can be 10 points. The specific value of the threshold B can be flexibly configured according to the actual application scenario.
[0221] 3. Locate AREA3 and perform ROI region detection based on YUV data2. For example, it can detect the regions to which faces, text, hair, game heroes, show characters, moving objects, etc. can be identified. The classification of these regions can be found in the description of the aforementioned embodiments, and will not be repeated here.
[0222] 4. For each type of AREA in AREA1-3, if there is overlap, they can be merged to obtain a larger AREA. Different types of AREAs will not be merged. For example, if there is overlap between two AREA1s, they will be merged. If there is overlap between an AREA1 and an AREA2, they will not be merged.
[0223] Then, the coordinate information of AREA1-3, the ROI region classification information, and VQA parameters are input into the code control and optimization module for processing:
[0224] In this method, as shown in Figure 12B, the quantization parameters corresponding to YUV data 2 are calculated and used as the original quantization parameters for AREA1-3 respectively. Then, for AREA1, QPOffset1 is subtracted from the original quantization parameters to obtain the optimized quantization parameters for AREA1, where the value range of QPOffset1 is [2,3]. For AREA2, QPOffset2 is subtracted from the original quantization parameters to obtain the optimized quantization parameters for AREA2. However, excessive adjustment of the quantization parameters in areas with large VQA fluctuations can make the overall video output bitrate uncontrollable; therefore, QPOffset1 is used to optimize the quantization parameters for AREA2. The value range of et2 can be [1,2]. For AREA3, based on the ROI classification information, QPOffset31 is subtracted from the original quantization parameters of the first type of ROI, QPOffset32 is subtracted from the original quantization parameters of the second type of ROI, and QPOffset33 is subtracted from the original quantization parameters of the third type of ROI. Here, QPOffset31 > QPOffset32 > QPOffset33. For example, QPOffset31 can be greater than or equal to 3, QPOffset32 can be greater than or equal to 2, and QPOffset33 can be greater than or equal to 1. It should be noted that the specific values of QPOffset1, QPOffset2, QPOffset31, QPOffset32, and QPOffset33 can be flexibly adjusted according to the application scenario, the allowable video bitrate fluctuation range, etc.
[0225] In the embodiment shown in Figure 14, relevant information from the image quality enhancement process is passed to the bitrate control process. This allows the bitrate control process to be combined with the image quality enhancement process, reducing the weakening effect of the bitrate control process on the image quality enhancement effect. This prevents some of the enhanced effects from being lost due to compression during the encoding process, thus improving the image quality of the encoded video frames and ultimately enhancing the user's Quality of Experience (QoE). Furthermore, applying this video encoding method to e-commerce live streaming can increase transaction volume by improving video quality. For example, see below...
[0226] As shown in Table 1, compared to not using the video encoding method shown in Figure 14, the number of orders in each live streaming room has increased positively.
[0227] Table 1
[0228] Referring to FIG15, FIG15 is a block diagram of a video encoding apparatus illustrating an exemplary embodiment of the present application. As shown in FIG15, the apparatus includes: an enhancement module 1501 configured to enhance the content of an original video frame contained in a video to be encoded, thereby obtaining an enhanced video frame; a processing module 1502 configured to detect the enhanced video frame, obtain content enhancement information of the enhanced video frame, and calculate encoding parameters of the enhanced video frame based on the content enhancement information; and an encoding module 1503 configured to encode the enhanced video frame according to the encoding parameters of the enhanced video frame, thereby obtaining encoded data of the enhanced video frame, and generate encoded data of the video to be encoded based on the encoded data of the enhanced video frame.
[0229] In an exemplary embodiment, based on the aforementioned scheme, the processing module 1502 is specifically configured to: select an optimization region to be optimized for encoding parameters from the enhanced video frame, and generate content enhancement information based on the optimization region; obtain the original encoding parameters of the optimization region, and optimize the original encoding parameters to obtain the optimized encoding parameters of the optimization region; wherein the image quality corresponding to the optimized encoding parameters is higher than the image quality corresponding to the original encoding parameters; and generate enhanced video frame encoding parameters based on the optimized encoding parameters of the optimization region.
[0230] In an exemplary embodiment, based on the aforementioned scheme, the processing module 1502 is specifically configured to: calculate the enhancement amplitude parameter of each region contained in the enhanced video frame, and find the region in the enhanced video frame whose corresponding enhancement amplitude parameter is greater than the enhancement amplitude threshold; and use the found region as the optimization region.
[0231] In an exemplary embodiment, based on the aforementioned scheme, and given that the enhancement amplitude threshold includes an image parameter threshold, the processing module 1502 is specifically configured to: segment the original video frame and the enhanced video frame according to the first size parameter, respectively, to obtain at least two original regions contained in the original video frame and at least two enhanced regions contained in the enhanced video frame; wherein, the enhanced regions correspond one-to-one with the original regions; calculate the image parameter change value corresponding to each enhanced region based on the image parameters of each enhanced region and the image parameters of the original region corresponding to each enhanced region, and use the image parameter change value as the enhancement amplitude parameter of each enhanced region; and search for the enhanced region whose corresponding image parameter change value is greater than the image parameter threshold from the at least two enhanced regions.
[0232] In an exemplary embodiment, based on the aforementioned scheme, and given that the enhancement amplitude threshold includes a video quality threshold, the processing module 1502 is specifically configured to: segment the enhanced video frame according to the second size parameter to obtain at least two candidate regions contained in the enhanced video frame; perform video quality detection on each candidate region to obtain video quality parameters for each candidate region, and calculate the enhancement amplitude parameter for each candidate region based on the video quality parameters of each candidate region; and search for candidate regions whose corresponding video quality parameters are greater than the video quality threshold from the at least two candidate regions.
[0233] In an exemplary embodiment, based on the aforementioned scheme, the processing module 1502 is specifically configured to: perform video quality detection on the enhanced video frame to obtain the video quality parameters of the enhanced video frame; and perform a difference calculation between the video quality parameters of each candidate region and the video quality parameters of the enhanced video frame to obtain the enhancement amplitude parameters of each candidate region.
[0234] In an exemplary embodiment, based on the aforementioned scheme, the processing module 1502 is specifically configured to: obtain video quality requirement information corresponding to the video to be encoded; calculate the enhancement amplitude threshold corresponding to the video to be encoded based on the video quality requirement information; wherein the enhancement amplitude threshold is negatively correlated with the video quality requirement information; and find the region in the enhanced video frame whose corresponding enhancement amplitude parameter is greater than the enhancement amplitude threshold corresponding to the video to be encoded.
[0235] In an exemplary embodiment, based on the aforementioned scheme, the processing module 1502 is specifically configured to: search for the target parameter optimization method corresponding to the optimized region from the parameter optimization methods corresponding to various region types; wherein the parameter optimization methods corresponding to different region types are not completely the same; optimize the original encoding parameters of the optimized region according to the target parameter optimization method to obtain the optimized encoding parameters of the optimized region.
[0236] In an exemplary embodiment, based on the aforementioned scheme, the processing module 1502 is specifically configured to: obtain the original quantization parameters of the optimized region from the original encoding parameters of the optimized region; reduce the original quantization parameters of the optimized region according to the target parameter optimization method to obtain the optimized quantization parameters of the optimized region; wherein the reduction range of the quantization parameters corresponding to different parameter optimization methods is not exactly the same; and generate the optimized encoding parameters of the optimized region based on the optimized quantization parameters of the optimized region.
[0237] In an exemplary embodiment, based on the aforementioned scheme, the processing module 1502 is specifically configured as follows: if the optimized region is a region of interest obtained by target detection of the enhanced video frame, then the target object to which the optimized region belongs is obtained; according to the importance of the target object, the original quantization parameter of the optimized region is reduced to obtain the optimized quantization parameter of the optimized region; wherein, the reduction of the original quantization parameter of the optimized region is positively correlated with the importance of the target object.
[0238] In an exemplary embodiment, based on the aforementioned scheme, the processing module 1502 is specifically configured to: obtain the bitrate control strategy corresponding to the video to be encoded, and calculate the encoding parameters corresponding to the enhanced video frame according to the bitrate control strategy; and calculate the original encoding parameters of the optimized region based on the encoding parameters corresponding to the enhanced video frame.
[0239] In an exemplary embodiment, based on the foregoing scheme, the processing module 1502 is further configured to: before obtaining the original encoding parameters of the optimized region and optimizing the original encoding parameters to obtain the optimized encoding parameters of the optimized region, search for optimized regions with overlapping regions from at least two optimized regions; merge the optimized regions with overlapping regions to obtain the merged optimized region.
[0240] It should be noted that the video encoding device provided in Figure 15 and the video encoding method provided in the above embodiments belong to the same concept. The specific way in which each module and unit performs operations has been described in detail in the method embodiments, and will not be repeated here.
[0241] Embodiments of this application also provide an electronic device, including: at least one processor; and a storage device for storing at least one computer program, which, when executed by the at least one processor, causes the electronic device to implement the video encoding methods provided in the above embodiments.
[0242] Figure 16 shows a schematic diagram of a computer system suitable for implementing an electronic device according to the embodiments of this application. The electronic device may be the terminal device 110 or the server 120 shown in Figure 1.
[0243] It should be noted that the computer system 1600 of the electronic device shown in Figure 16 is only an example and should not impose any limitations on the functionality and scope of use of the embodiments of this application.
[0244] As shown in Figure 16, the computer system 1600 includes a Central Processing Unit (CPU) 1601, which can perform various appropriate actions and processes based on computer programs stored in Read-Only Memory (ROM) 1602 or loaded from storage portion 1608 into Random Access Memory (RAM) 1603, such as executing the video encoding method described in the above embodiments. The RAM 1603 also stores various computer programs and data required for system operation. The CPU 1601, ROM 1602, and RAM 1603 are interconnected via a bus 1604. An Input / Output (I / O) interface 1605 is also connected to the bus 1604.
[0245] In some embodiments, the following components are connected to the I / O interface 1605: an input section 1606 including a keyboard, mouse, etc.; an output section 1607 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 1608 including a hard disk, etc.; and a communication section 1609 including a network interface card such as a LAN (Local Area Network) card, modem, etc. The communication section 1609 performs communication processing via a network such as the Internet. A drive 1610 is also connected to the I / O interface 1605 as needed. A removable medium 1611, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed on the drive 1610 as needed so that computer programs read from it can be installed into the storage section 1608 as needed.
[0246] In particular, according to embodiments of this application, a computer program implementing the video encoding method can be carried on a computer-readable medium, which can be downloaded and installed from a network via the communication section 1609, and / or installed from a removable medium 1611.
[0247] It should be noted that the computer-readable medium shown in the embodiments of this application can be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. For example, a computer-readable storage medium can be an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or at least two wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), flash memory, optical fiber, portable compact disc read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this application, a computer-readable storage medium can be any tangible medium containing or storing a computer program that can be used by or in conjunction with an instruction execution system, apparatus, or device. Computer-readable signal media may include data signals propagated in baseband or as part of a carrier wave, carrying a computer-readable computer program. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. The computer program contained in the computer-readable medium may be transmitted using any suitable medium, including but not limited to wireless, wired, etc., or any suitable combination thereof.
[0248] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. Each block in a flowchart or block diagram may represent a module, segment, or portion of code, which contains one or at least two executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and a computer program.
[0249] The units described in the embodiments of this application can be implemented in software or hardware, and the described units can also be located in a processor. The names of these units do not necessarily limit the specific unit itself.
[0250] Another aspect of this application provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor of an electronic device, causes the electronic device to implement the aforementioned video encoding method. This computer-readable storage medium may be included in the electronic device described in the above embodiments, or it may exist independently and not assembled into the electronic device.
[0251] Another aspect of this application provides a computer program product, which includes a computer program that, when executed by a processor, implements the video encoding methods provided in the various embodiments described above. The computer program can be stored in a computer-readable storage medium. The computer program product can be a computer program as a product, such as an APP (Application), webpage, mini-program, etc.; or, the computer program product can also be a storage medium, device, terminal, virtual machine, etc., containing the computer program.
[0252] Another aspect of this application provides a method for processing video streams generated according to the video encoding method described in the foregoing embodiments. The encoded data of the video to be encoded is the basis and source for video stream generation. After obtaining the encoded data of the enhanced video frames, the encoded data of the video to be encoded is then generated based on this encoded data. This encoded data is organized according to specific formats and rules, ultimately forming a video stream that can be used for storage, playback, and transmission.
[0253] In practical applications, video stream processing can be applied to several key scenarios. In storage scenarios, video streams need to be encapsulated according to specific file formats to achieve long-term, stable preservation and facilitate subsequent retrieval. For example, common video file formats such as MP4 and AVI have their own defined stream organization methods, which helps achieve compatibility and efficient storage across different devices and systems. In playback scenarios, the core of video stream processing lies in decoding the encoded data to restore directly displayable video frames. To improve the user's viewing experience, a series of post-processing operations may be performed on the decoded video frames, such as noise reduction to reduce noise interference and make the image clearer; and image quality enhancement by adjusting parameters such as color and contrast to make the picture more vivid and realistic. In transmission scenarios, to reduce bandwidth consumption, improve transmission efficiency, and ensure data security during transmission, video streams are typically compressed and encrypted. For example, advanced compression algorithms are used to compress the stream, reducing data volume while maintaining video quality as much as possible; encryption technology is used to encrypt the stream to prevent data from being stolen or tampered with during transmission.
[0254] Another aspect of this application provides a computer-readable storage medium having computer instructions stored thereon. When the computer instructions are executed by at least one processor, they implement the video encoding method described in the foregoing embodiments, generate a video stream based on the encoded data of the video, and store the video stream.
[0255] In summary, this application provides a video encoding method, apparatus, electronic device, computer-readable storage medium, computer program product, and method for processing video bitstreams. The method involves enhancing the content of original video frames in the video to be encoded to obtain enhanced video frames; detecting the enhanced video frames to obtain content enhancement information; calculating encoding parameters for the enhanced video frames based on the content enhancement information; encoding the enhanced video frames according to the encoding parameters to obtain encoded data for the enhanced video frames; and generating encoded data for the video to be encoded based on the encoded data of the enhanced video frames. By performing content enhancement before encoding the video frames, the image quality of the encoded video frames is improved. Furthermore, since the encoding parameters are calculated based on the content enhancement information, the content enhancement process is deeply integrated with the encoding process, reducing the weakening effect of the enhancement process during encoding, minimizing the loss of video frame details, and improving the quality and visual effect of the encoded video.
[0256] Furthermore, when calculating the coding parameters of the enhanced video frame, an optimization region to be optimized is selected from the enhanced video frame. Based on the optimization region, content enhancement information is generated, the original coding parameters of the optimization region are obtained, and the original coding parameters are optimized to obtain the optimized coding parameters of the optimization region. Based on the optimized coding parameters of the optimization region, the coding parameters of the enhanced video frame are generated. This ensures the image quality, such as the sharpness, after encoding the optimization region, improves the image quality of the optimized region after encoding, and reduces the weakening effect of the encoding process on the enhancement effect.
[0257] Furthermore, when selecting optimization regions, the enhancement magnitude parameter of each region contained in the enhanced video frame is calculated, and regions with enhancement magnitude parameters greater than the enhancement magnitude threshold are found in the enhanced video frames. These found regions are then selected as optimization regions. Selecting regions with larger enhancement magnitudes as optimization regions avoids the loss of some enhancement effects obtained from content enhancement during encoding and compression, thereby improving the image quality of the encoded video frames.
[0258] Furthermore, when the enhancement amplitude threshold includes an image parameter threshold, the original video frame and the enhanced video frame are segmented according to the first size parameter, resulting in at least two original regions contained in the original video frame and at least two enhanced regions contained in the enhanced video frame. Based on the image parameters of each enhanced region and the image parameters of the corresponding original region, the change value of the image parameters for each enhanced region is calculated, and this change value is used as the enhancement amplitude parameter for each enhanced region. The enhanced region whose corresponding image parameter change value is greater than the image parameter threshold is then searched among the at least two enhanced regions. By comparing the image parameters of the original region and the enhanced region, regions with larger enhancement amplitudes can be accurately identified, providing more precise region selection for subsequent optimization.
[0259] Furthermore, when the enhancement amplitude threshold includes a video quality threshold, the enhanced video frame is segmented according to the second size parameter to obtain at least two candidate regions contained in the enhanced video frame. Video quality detection is performed on each candidate region to obtain the video quality parameter of each candidate region. Based on the video quality parameter of each candidate region, the enhancement amplitude parameter of each candidate region is calculated. Candidate regions with video quality parameters greater than the video quality threshold are then searched from the at least two candidate regions. Determining the enhancement amplitude parameter based on video quality detection can filter out regions that need optimization from the perspective of video quality, ensuring the encoding effect of high-quality regions in the video.
[0260] Furthermore, when calculating the enhancement magnitude parameter for each candidate region based on its video quality parameters, video quality detection is performed on the enhanced video frame to obtain its video quality parameters. The difference between the video quality parameters of each candidate region and those of the enhanced video frame is then calculated to obtain the enhancement magnitude parameter for each candidate region. This difference calculation more accurately reflects the enhancement magnitude of each candidate region relative to the overall enhanced video frame, providing a more scientific basis for region selection.
[0261] Furthermore, when searching for regions where the enhancement amplitude parameter is greater than the enhancement amplitude threshold, the video quality requirement information corresponding to the video to be encoded is obtained. Based on the video quality requirement information, the enhancement amplitude threshold corresponding to the video to be encoded is calculated. The enhancement amplitude threshold is negatively correlated with the video quality requirement information. Regions where the enhancement amplitude parameter is greater than the enhancement amplitude threshold corresponding to the video to be encoded are then searched from the enhanced video frames. Dynamically adjusting the enhancement amplitude threshold according to the video quality requirements can rationally allocate computing resources while meeting different video quality needs, thereby improving video encoding efficiency.
[0262] Furthermore, when optimizing the original encoding parameters, the target parameter optimization method for the optimized region is selected from the parameter optimization methods corresponding to various region types. The parameter optimization methods for different region types are not entirely the same. Based on the target parameter optimization method, the original encoding parameters of the optimized region are optimized to obtain the optimized encoding parameters for the optimized region. Using different parameter optimization methods for different types of optimized regions can meet different business needs and improve the flexibility of parameter optimization.
[0263] Furthermore, when optimizing the original coding parameters of the optimization region according to the target parameter optimization method, the original quantization parameters of the optimization region are obtained from the original coding parameters of the optimization region. Based on the target parameter optimization method, the original quantization parameters of the optimization region are reduced to obtain the optimized quantization parameters of the optimization region. The reduction in quantization parameters varies depending on the optimization method. Based on the optimized quantization parameters of the optimization region, the optimized coding parameters of the optimization region are generated. Directly reducing the original quantization parameters to optimize the original coding parameters reduces the complexity of parameter optimization, saves computational resources, and improves video coding efficiency.
[0264] Furthermore, when reducing the original quantization parameters of the optimization region, if the optimization region is a region of interest obtained by object detection of the enhanced video frame, the target object to which the optimization region belongs is obtained. Based on the importance of the target object, the original quantization parameters of the optimization region are reduced to obtain the optimized quantization parameters. The reduction in the original quantization parameters of the optimization region is positively correlated with the importance of the target object. Adjusting the reduction in quantization parameters according to the importance of the target object can improve the image quality of the video frame while reasonably controlling the consumption of computational resources, thus improving the efficiency of video encoding.
[0265] Furthermore, when obtaining the original encoding parameters of the optimization region, the bitrate control strategy corresponding to the video to be encoded is obtained. Based on the bitrate control strategy, the encoding parameters corresponding to the enhanced video frames are calculated. Based on the encoding parameters corresponding to the enhanced video frames, the original encoding parameters of the optimization region are calculated. By deeply integrating the content enhancement process with the bitrate control process, the impact of bitrate control on the enhancement effect is reduced while ensuring the bitrate control effect, thus improving the enhancement effect and image quality of the encoded video.
[0266] Furthermore, when there are at least two optimization regions, before obtaining the original coding parameters of the optimization regions and optimizing them to obtain the optimized coding parameters, overlapping optimization regions are identified from the at least two optimization regions. These overlapping optimization regions are then merged to obtain the merged optimization region. Merging overlapping optimization regions saves computational resources and improves video coding efficiency.
[0267] Furthermore, when calculating the encoding parameters for enhanced video frames based on different bitrate control strategies (such as Constant Quantization Parameter (CQP), Constant Quality Factor (CRF), Specified Average Bitrate (ABR), Constant Bitrate (CBR), Video Buffer Verification (VBV), and Two-Pass Encoding), different strategies can meet the diverse application scenario requirements. For example, the Constant Quantization Parameter strategy ensures constant quantization distortion for each video frame, making the encoding process more stable; the Constant Quality Factor strategy effectively balances quality and bitrate; the Specified Average Bitrate strategy controls bandwidth usage; the Constant Bitrate strategy ensures stable data volume per unit time; the Video Buffer Verification strategy avoids stuttering during video playback; and the Two-Pass Encoding strategy analyzes the video content in the first pass and performs more refined encoding based on the analysis results in the second pass, thereby improving encoding efficiency and quality. Through these different strategies, encoding resources can be rationally allocated in different scenarios, improving encoding efficiency and video quality.
[0268] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0269] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this patent application should be determined by the appended claims.
Claims
1. A video encoding method, performed by an electronic device, the method comprising: Content enhancement is performed on the original video frames contained in the video to be encoded to obtain enhanced video frames; The enhanced video frame is detected to obtain the content enhancement information of the enhanced video frame, and the encoding parameters of the enhanced video frame are calculated based on the content enhancement information; and The enhanced video frame is encoded according to the enhanced video frame encoding parameters to obtain the encoded data of the enhanced video frame, and the encoded data of the video to be encoded is generated based on the encoded data of the enhanced video frame.
2. The method as described in claim 1, wherein detecting the enhanced video frame to obtain content enhancement information of the enhanced video frame, and calculating the encoding parameters of the enhanced video frame based on the content enhancement information, includes: The optimization region to be optimized for encoding parameters is selected from the enhanced video frame, and the content enhancement information is generated based on the optimization region; The original coding parameters of the optimized region are obtained, and the original coding parameters are optimized to obtain the optimized coding parameters of the optimized region; wherein the image quality corresponding to the optimized coding parameters is higher than the image quality corresponding to the original coding parameters. The enhanced video frame coding parameters are generated based on the optimized coding parameters of the optimized region.
3. The method as described in claim 2, wherein selecting the optimization region to be optimized for encoding parameters from the enhanced video frame includes: Calculate the enhancement amplitude parameter of each region contained in the enhanced video frame, and find the region in the enhanced video frame whose corresponding enhancement amplitude parameter is greater than the enhancement amplitude threshold; The identified region is used as the optimization region.
4. The method of claim 3, wherein the enhancement amplitude threshold includes an image parameter threshold; The step of calculating the enhancement magnitude parameter of each region contained in the enhanced video frame, and finding the region in the enhanced video frame whose corresponding enhancement magnitude parameter is greater than the enhancement magnitude threshold, includes: According to the first size parameter, the original video frame and the enhanced video frame are respectively segmented to obtain at least two original regions contained in the original video frame and at least two enhanced regions contained in the enhanced video frame; wherein, the enhanced regions correspond one-to-one with the original regions; Based on the image parameters of each enhanced region and the image parameters of the original region corresponding to each enhanced region, the change value of the image parameters corresponding to each enhanced region is calculated, and the change value of the image parameters is used as the enhancement amplitude parameter of each enhanced region. From the at least two enhancement regions, find the enhancement region whose corresponding image parameter change value is greater than the image parameter threshold.
5. The method of claim 3, wherein the enhancement amplitude threshold includes a video quality threshold; The step of calculating the enhancement magnitude parameter of each region contained in the enhanced video frame, and finding the region in the enhanced video frame whose corresponding enhancement magnitude parameter is greater than the enhancement magnitude threshold, includes: The enhanced video frame is segmented according to the second size parameter to obtain at least two candidate regions contained in the enhanced video frame; For each candidate region, video quality detection is performed to obtain the video quality parameters of each candidate region, and the enhancement amplitude parameter of each candidate region is calculated based on the video quality parameters of each candidate region. From the at least two candidate regions, find the candidate regions whose corresponding video quality parameters are greater than the video quality threshold.
6. The method of claim 5, wherein calculating the enhancement magnitude parameter of each candidate region based on the video quality parameters of each candidate region includes: The enhanced video frame is subjected to video quality detection to obtain the video quality parameters of the enhanced video frame; The difference between the video quality parameters of each candidate region and the video quality parameters of the enhanced video frame is calculated to obtain the enhancement amplitude parameter of each candidate region.
7. The method according to any one of claims 3 to 6, wherein finding the region in the enhanced video frame whose corresponding enhancement amplitude parameter is greater than the enhancement amplitude threshold comprises: Obtain the video quality requirement information corresponding to the video to be encoded; Based on the video quality requirement information, calculate the enhancement threshold corresponding to the video to be encoded; wherein, the enhancement threshold is negatively correlated with the video quality requirement information; Find the region in the enhanced video frame whose enhancement amplitude parameter is greater than the enhancement amplitude threshold corresponding to the video to be encoded.
8. The method according to any one of claims 2 to 7, wherein optimizing the original coding parameters to obtain optimized coding parameters for the optimized region comprises: From the parameter optimization methods corresponding to various region types, find the target parameter optimization method corresponding to the optimized region; wherein, the parameter optimization methods corresponding to different region types are not completely the same; Based on the target parameter optimization method, the original encoding parameters of the optimization region are optimized to obtain the optimized encoding parameters of the optimization region.
9. The method as described in claim 8, wherein optimizing the original coding parameters of the optimization region according to the target parameter optimization method to obtain the optimized coding parameters of the optimization region includes: Obtain the original quantization parameters of the optimized region from the original encoding parameters of the optimized region; Based on the target parameter optimization method, the original quantization parameter of the optimization region is reduced to obtain the optimized quantization parameter of the optimization region; wherein, the reduction in quantization parameter corresponding to different parameter optimization methods is not exactly the same. Based on the optimized quantization parameters of the optimized region, optimized coding parameters for the optimized region are generated.
10. The method of claim 9, wherein reducing the original quantization parameter of the optimization region according to the target parameter optimization method to obtain the optimized quantization parameter of the optimization region includes: If the optimized region is a region of interest obtained by target detection of the enhanced video frame, then the target object to which the optimized region belongs is obtained; Based on the importance of the target object, the original quantization parameter of the optimization region is reduced to obtain the optimized quantization parameter of the optimization region; wherein, the reduction in the original quantization parameter of the optimization region is positively correlated with the importance of the target object.
11. The method according to any one of claims 2 to 10, wherein obtaining the original encoding parameters of the optimized region comprises: Obtain the bitrate control strategy corresponding to the video to be encoded, and calculate the encoding parameters corresponding to the enhanced video frame based on the bitrate control strategy; Based on the encoding parameters corresponding to the enhanced video frame, the original encoding parameters of the optimized region are calculated.
12. The method of any one of claims 2 to 11, wherein the number of optimized regions comprises at least two; Before obtaining the original coding parameters of the optimized region and optimizing the original coding parameters to obtain the optimized coding parameters of the optimized region, the method further includes: Find the optimization regions that overlap between at least two optimization regions; The optimization regions with overlapping areas are merged to obtain the merged optimization region.
13. A video encoding apparatus, the apparatus comprising: The enhancement module is configured to enhance the content of the original video frames contained in the video to be encoded, thereby obtaining enhanced video frames; The processing module is configured to detect the enhanced video frame, obtain the content enhancement information of the enhanced video frame, and calculate the encoding parameters of the enhanced video frame based on the content enhancement information; and The encoding module is configured to encode the enhanced video frame according to the enhanced video frame encoding parameters to obtain the encoded data of the enhanced video frame, and generate the encoded data of the video to be encoded based on the encoded data of the enhanced video frame.
14. An electronic device comprising: At least one processor; A storage device for storing at least one computer program, which, when executed by the at least one processor, causes the electronic device to implement the video encoding method according to any one of claims 1-12.
15. A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor of an electronic device, causes the electronic device to implement the video encoding method according to any one of claims 1-12.
16. A computer program product comprising a computer program that, when executed by a processor, implements the video encoding method according to any one of claims 1-12.
17. A method for processing a video stream, said video stream being generated by the video encoding method according to any one of claims 1-12.
18. A computer-readable storage medium having stored thereon computer instructions that, when executed by at least one processor, implement the video encoding method of any one of claims 1-12 to generate and store a video stream.