Video frame coding method and device, equipment and medium
By identifying sensitive areas in video frames and performing differentiated bitrate allocation and quantization parameter adjustment, the problem of mosaic distortion in low bitrate video encoding is solved, improving video quality and subjective viewing experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING BAIDU NETCOM SCI & TECH CO LTD
- Filing Date
- 2026-01-28
- Publication Date
- 2026-05-01
AI Technical Summary
Existing low bitrate video encoding technologies are prone to mosaic distortion when the compression ratio is increased, especially affecting the subjective viewing experience in sensitive areas. Furthermore, existing optimization strategies lack specificity and precision, resulting in an uneven bitrate distribution.
By identifying sensitive areas in video frames, such as smoke, sky, skin, and lights, and combining the area ratio, visual sensitivity, and texture complexity, differentiated bitrate allocation and quantization parameter adjustment are performed, and then encoded using a video encoder.
It effectively reduces the impact of mosaic distortion, improves the image quality of areas sensitive to the human eye in video frames, and avoids the risk of wasted bitrate and overall bitrate exceeding budget.
Smart Images

Figure CN121967686A_ABST
Abstract
Description
A video frame encoding method, apparatus, device and medium Technical Field
[0001] This disclosure relates to the fields of artificial intelligence and cloud computing, specifically to image processing and video encoding scenarios. Background Technology
[0002] With the rapid development of 5G and IoT technologies, the demand for video data transmission and storage is increasing daily. Low-bitrate video encoding technology, due to its ability to effectively reduce transmission bandwidth and storage costs, has become a core requirement in many scenarios, such as video surveillance in remote areas, short video transmission on mobile devices, and video backhaul in satellite communications. In low-bitrate scenarios, video encoders need to control the bitrate by increasing the compression ratio. However, high compression ratios often lead to video distortion, among which mosaic distortion is one of the most common and has a significant impact on subjective viewing experience. How to reduce the impact of mosaic distortion on subjective experience is an important issue in the industry. Summary of the Invention
[0003] This disclosure provides a video frame encoding method, apparatus, device, and medium.
[0004] According to one aspect of this disclosure, a video frame encoding method is provided, the method comprising: identifying sensitive regions in a video frame to obtain target regions; allocating bitrate to the target regions based on a basic encoding strategy, the overall base bitrate of the video frame, and the area ratio, visual sensitivity, and texture complexity of the target regions to obtain bitrate allocation results for the target regions; adjusting quantization parameters of the target regions and non-target regions in the video frame based on the bitrate allocation results to obtain differentiated encoding parameters; and encoding the video frame using a video encoder according to the differentiated encoding parameters.
[0005] According to another aspect of this disclosure, a video frame encoding apparatus is provided, comprising: a target region determination module for identifying sensitive regions in a video frame to obtain a target region; a bitrate allocation module for allocating bitrate to the target region based on a basic encoding strategy, the overall basic bitrate of the video frame, and the area ratio, visual sensitivity, and texture complexity of the target region, to obtain a bitrate allocation result for the target region; a parameter adjustment module for adjusting quantization parameters of the target region and non-target regions in the video frame based on the bitrate allocation result, to obtain differentiated encoding parameters; and a video encoding module for encoding the video frame using a video encoder according to the differentiated encoding parameters.
[0006] According to another aspect of this disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the video frame encoding method according to any embodiment of this disclosure.
[0007] According to another aspect of this disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to cause a computer to perform the video frame encoding method described in any embodiment of this disclosure.
[0008] According to another aspect of this disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements the video frame encoding method described in any embodiment of this disclosure.
[0009] The technology disclosed herein can improve the quality of video images.
[0010] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0011] The accompanying drawings are provided to better understand the present solution and do not constitute a limitation of the present disclosure. Specifically: Figure 1 is a flowchart of a video frame encoding method according to an embodiment of the present disclosure; Figure 2 is a flowchart of another video frame encoding method according to an embodiment of the present disclosure; Figure 3 is a structural schematic diagram of a video frame encoding apparatus according to an embodiment of the present disclosure; and Figure 4 is a block diagram of an electronic device used to implement the video frame encoding method of the embodiments of the present disclosure. Detailed Implementation
[0012] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.
[0013] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0014] Furthermore, it should be noted that the collection, storage, use, processing, transmission, provision, and disclosure of video frame-related data involved in the technical solution of this invention all comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0015] Mosaic distortion is essentially caused by insufficient bitrate during video encoding, leading to inaccurate block matching and excessive quantization errors. This results in a noticeable blocky effect at image block boundaries, presenting a mosaic-like visual effect. Although existing video coding standards (such as H.264, H.265 / HEVC, and H.266 / VVC) employ a series of optimization techniques, including intra-frame prediction, inter-frame prediction, transform coding, entropy coding, and post-processing, mosaic distortion is still difficult to avoid in low-bitrate compression scenarios.
[0016] Further research revealed that the impact of mosaic distortion on subjective experience exhibits significant regional differences: the human eye is more sensitive to mosaic distortion in specific subjectively sensitive areas such as smoke, sky, skin, and lighting. For example, the skin area, as the core area of a person in a video, is quickly detected by the human eye even with slight mosaic distortion, severely impacting the viewing experience; areas such as the sky and smoke, due to their smooth textures, suffer from mosaic distortion that disrupts visual continuity, creating a noticeable blocky segmentation; and lighting areas, with their high brightness and strong contrast, experience uneven halo diffusion due to mosaic distortion, further exacerbating subjective discomfort.
[0017] Existing technologies for balancing bitrate and image quality mainly employ the following optimization schemes: 1) Bitrate control optimization: Strategies such as Constant Bit Rate (CBR) and Variable Bit Rate (VBR) are used to dynamically adjust the quantization parameters of the current frame by statistically analyzing the encoding information of the preceding frames (such as the number of encoded bits and distortion) to avoid excessive fluctuations in the bitrate of a single frame; for example, the bitrate control algorithm of H.265 / HEVC predicts the bitrate requirement of the current frame through a linear model and dynamically corrects the quantization parameter (QP) value to match the overall bitrate budget. 2) Spatial Adaptive Quantization (AQ) technology: As a core branch of quantization optimization, its core implementation logic is based on the spatial masking effect of the human visual system, dynamically allocating quantization resources to different spatial locations of the image; 3) Loop Filter Optimization: Modules such as De-blocking Filtering (DBK), Sample Adaptive Offset (SAO), and Adaptive Loop Filter (ALF) are added to smooth the boundaries of encoded image blocks and alleviate block distortion; 4) Prediction Mode Optimization: The number of intra-frame prediction modes is expanded (e.g., H.265 / HEVC provides 35 intra-frame prediction modes, and Versatile Video Coding (VVC) expands to 67 modes), improving the prediction accuracy of flat and textured regions, reducing prediction residuals, and thus reducing bitrate requirements.
[0018] Existing low bitrate encoding optimization techniques suffer from the following main drawbacks: 1) Lack of targeted bitrate allocation: Existing techniques mostly employ a global bitrate allocation strategy, distributing the bitrate evenly across the entire video frame without considering the differences in human eye sensitivity to different regions. Under low bitrate budgets, sensitive areas prone to pixelation cannot receive sufficient bitrate support, leading to severe distortion; while areas less sensitive to human vision (such as complex textured backgrounds) are allocated too much bitrate, resulting in wasted bitrate. 2) Insufficient accuracy in sensitive area recognition: Existing AQ (Advanced Quality Qualification) techniques' region segmentation and weighting models are mostly based on general texture features, without customized designs for specific subjectively sensitive areas such as smoke, sky, and skin. They also lack semantic information, making it difficult to accurately identify specific types of sensitive areas such as smoke, sky, skin, and lighting. For example, they may misclassify smooth-textured background areas as sensitive areas or miss low-contrast skin areas, resulting in poor optimization performance. 3) Single encoding optimization strategy: Existing optimization methods mostly improve the quality of local areas by fixing and reducing quantization parameters, without dynamically adjusting the optimization intensity based on the sensitivity level and texture complexity of the region. This approach either fails to solve the mosaic problem due to insufficient optimization, or it causes the overall bitrate to exceed the budget due to over-optimization, thus losing the advantage of low bitrate encoding.
[0019] Figure 1 is a flowchart of a video frame encoding method according to an embodiment of the present disclosure; this method is applicable to situations where mosaic distortion can be avoided in low bitrate extreme compression scenarios. This method can be executed by a video frame encoding device, which can be implemented in software and / or hardware and integrated into an electronic device carrying video frame encoding functionality, such as a server. As shown in Figure 1, the video frame encoding method of this embodiment may include: S101, identifying sensitive regions of the video frame to obtain the target region.
[0020] In this embodiment, the target area refers to the area that is sensitive to the human eye and prone to pixelation, such as smoke, sky, skin, and light.
[0021] One alternative approach is to identify sensitive regions in video frames based on a semantic segmentation model to obtain the target region. Here, the semantic segmentation model refers to a pre-trained U-Net semantic segmentation network.
[0022] S102, based on the basic coding strategy, the overall base bitrate of the video frame, and the area ratio, visual sensitivity, and texture complexity of the target region, bitrate allocation is performed for the target region to obtain the bitrate allocation result of the target region.
[0023] In this embodiment, the basic coding strategy refers to the conventional coding strategy for video encoding. The overall base bitrate refers to the base allocated bitrate of the video frame. The region area ratio refers to the proportion of the target region's area within the total area of the video frame. Visual sensitivity is used to evaluate the sensitivity of the candidate region to the human eye. Texture complexity is used to assess the complexity of the texture features of the initial region. The bitrate allocation result refers to the bitrate allocated to the target region. The bitrate allocation result includes the target bitrate allocated to the target region.
[0024] An alternative approach involves allocating bitrate to a target region based on a bitrate allocation model, a basic coding strategy, the overall base bitrate of the video frames, the area ratio of the target region, visual sensitivity, and texture complexity, thus obtaining the bitrate allocation result for the target region. The bitrate allocation model can be obtained by training a neural network model based on the basic coding strategy and the bitrate allocation of the video after encoding.
[0025] S103, based on the bitrate allocation result, adjust the quantization parameters of the target region and non-target region in the video frame to obtain differentiated coding parameters.
[0026] In this embodiment, the differential coding parameters refer to the coding parameters of different regions in a video frame.
[0027] Specifically, based on the bitrate allocation results, differentiated encoding parameters are applied to target and non-target regions to achieve targeted optimization. The core optimization strategy is the adaptive adjustment of the quantization parameter (QP). The quantization parameter directly determines the degree of encoding distortion; the smaller the quantization parameter value, the less distortion, but the higher the bitrate. Therefore, the following optimization methods are used for target and non-target regions to obtain differentiated encoding parameters: For the target region, a direct mapping relationship between the target bitrate and the quantization parameter is established. Based on this direct mapping relationship, the first encoding parameter for the target region is determined. This ensures that the adjusted QP accurately matches the dynamically allocated target bitrate. The formula corresponding to the mapping relationship is shown below: Among them, The standard QP value for non-target regions. Quantization parameters for the adjusted target region. Bitrate-QP sensitivity coefficient; Indicates the reference bitrate for the target region; The total bitrate of the target region is obtained by summing the reference bitrate and the additional bitrate of the target region.
[0028] For non-target regions, conventional encoding parameters are used to obtain the second encoding parameters for the non-target regions, ensuring no significant distortion under the basic bitrate allocation while avoiding bitrate waste.
[0029] The first encoding parameter of the target region and the second encoding parameter of the non-target region are used as the differential encoding adoption number.
[0030] Furthermore, for the target area, intra-frame prediction mode optimization can be used to increase the number of intra-frame prediction modes in the target area, making the number of modes more accurate, improving block matching accuracy, and further reducing mosaic distortion.
[0031] It should be noted that by adjusting the adaptive quantization parameters and optimizing intra-frame prediction, mosaic distortion in the target region is effectively suppressed.
[0032] S104 uses a video encoder to encode video frames according to differentiated encoding parameters.
[0033] Specifically, the differential coding parameters corresponding to each video frame are input into the video encoder. The encoder encodes the video frame according to the differential coding parameters and outputs the encoded video stream.
[0034] It should be noted that the encoding format uses H.256, which has 35 variations, and the encoding format of the video frames can be flexibly adjusted.
[0035] The technical solution provided in this disclosure identifies target regions by performing sensitive region identification on video frames. Based on a basic encoding strategy, the overall base bitrate of the video frame, and the area ratio, visual sensitivity, and texture complexity of the target region, bitrate allocation is performed on the target region to obtain a bitrate allocation result. Based on the bitrate allocation result, quantization parameters are adjusted for target and non-target regions in the video frame to obtain differentiated encoding parameters. A video encoder is then used to encode the video frame according to these differentiated encoding parameters. This technical solution, by combining the overall base bitrate of the video frame with the area ratio, visual sensitivity, and texture complexity of the target region for bitrate allocation, and then performing differentiated encoding on target and non-target regions in the video frame, can improve the image quality of visually sensitive mosaic areas in video frames.
[0036] Based on the above embodiments, as an optional approach of the present invention, sensitive region identification is performed on video frames to obtain target regions, including: semantic segmentation of video frames based on preset sensitive categories to obtain initial regions belonging to sensitive categories; calculating the texture complexity of the initial regions and filtering candidate regions prone to mosaic based on the texture complexity from the initial regions; determining the visual sensitivity of the candidate regions and filtering target regions from the candidate regions based on the visual sensitivity.
[0037] Sensitive categories include smoke, sky, skin, and light. The initial region refers to the location range of sensitive areas preliminarily determined through semantic segmentation. Candidate regions refer to areas prone to pixelation selected from the initial region based on texture complexity.
[0038] Specifically, firstly, an improved lightweight U-Net semantic segmentation model is used to perform semantic segmentation on video frames, outputting a semantic segmentation mask. Each pixel in the semantic segmentation mask corresponds to a predefined sensitive category, i.e., a semantic category label, thus obtaining the initial region belonging to the sensitive category. The improved lightweight U-Net model replaces the standard convolution in traditional U-Net with depthwise separable convolutions, reducing the number of network parameters and computational cost while maintaining segmentation accuracy, thereby reducing computational complexity.
[0039] Then, for each initial region, the texture complexity of that region is determined. For example, the texture features of the initial region can be extracted using a gray-level co-occurrence matrix, and the texture contrast and entropy value can be calculated based on these features. These texture contrast and entropy values are then used as the texture complexity. Subsequently, initial regions with texture contrast less than or equal to a contrast threshold and entropy values less than or equal to an entropy threshold are selected as candidate regions for mosaic. The contrast threshold and entropy threshold can be adjusted based on the actual scene requirements; for example, the contrast threshold can range from 0.03 to 0.15, and the entropy threshold can range from 1.0 to 2.0. Texture contrast reflects the degree of difference in pixel gray levels within the initial region; regions with smooth textures have lower contrast. The texture contrast of the initial region is determined using the following formula. : Entropy reflects the complexity of the initial texture region; regions with smooth textures have lower entropy values. The entropy value can be determined using the following formula. : .
[0040] in It is the minimum value; This represents the square of the difference in standard grayscale values. This indicates that the grayscale values i and j are at a specified distance. With angle The probability of coexistence.
[0041] Next, human visual sensitivity is closely related to the brightness, contrast, and color features of a region: regions with moderate brightness are more sensitive than excessively bright or dark regions; regions with higher contrast are more sensitive; and natural color regions such as skin tone and sky blue are more sensitive than other color regions. Therefore, for each candidate region, its brightness features, texture contrast, and color histogram features are extracted and input into the visual sensitivity model. The model outputs the visual sensitivity of the candidate region. The visual sensitivity model is constructed based on the characteristics of the human visual system, and its output is the visual sensitivity V of the candidate region. The model is specifically expressed by the following formula: ;in, To represent brightness features, the average pixel value of the Y channel in the YUV format of the candidate region is extracted, and the Y channel directly represents brightness information. Indicates texture contrast; Represents the characteristics of the color histogram. The smaller the value, the more uniform the color of the area, the closer it is to a pure color area, and the higher the sensitivity of the human eye. Color characteristics are represented by calculating the mean and variance of the U and V channels. It should be noted that the higher the visual sensitivity, the stronger the sensitivity; the value range of visual sensitivity is 0-10. , and These are the fusion weights, which can be set according to the actual situation, with 0.4, 0.3, and 0.3 being preferred.
[0042] Finally, candidate regions with visual sensitivity greater than or equal to the sensitivity threshold are selected as target regions. The sensitivity threshold can be adjusted according to the actual needs of the scenario.
[0043] Understandably, combining semantic segmentation, texture complexity, and visual sensitivity to identify target areas, i.e., mosaic areas in the human eye, can not only accurately locate specific types of sensitive areas such as smoke, sky, skin, and light, but also eliminate non-mosaic areas and low-sensitivity areas through texture complexity filtering and sensitivity scoring, thus accurately identifying target areas, i.e. mosaic areas, providing a precise regional basis for targeted optimization.
[0044] Based on the above embodiments, as an optional approach of this disclosure, the input original video frame is preprocessed to obtain a video frame. Specifically, this can be done by using a Gaussian filter to denoise the original video frame, balancing the denoising effect with detail preservation, then scaling the denoised video frame to a preset size such as 854×480, 640×360, etc., using a bilinear interpolation algorithm to ensure the smoothness of the scaled video frame, and finally using a normalization formula to normalize the scaled video frame, mapping the pixel values to the range [0,1], thereby improving the convergence speed and recognition accuracy of semantic seams; wherein, the normalization formula is as follows: ;in, Represents the original pixel value. and These represent the maximum and minimum pixel values of the video frame, respectively.
[0045] Understandably, preprocessing the original video frames can reduce the interference of noise on subsequent region recognition and encoding optimization, while also unifying the input size to adapt to subsequent network models.
[0046] Figure 2 is a flowchart of another video frame encoding method provided according to an embodiment of this disclosure. Based on the above embodiments, this embodiment further optimizes the process of "allocating bitrate to the target region according to the basic encoding strategy, the overall base bitrate of the video frame, and the area ratio, visual sensitivity, and texture complexity of the target region, to obtain the bitrate allocation result of the target region," providing an optional implementation scheme. As shown in Figure 2, the method includes: S201, identifying sensitive regions of the video frame to obtain the target region.
[0047] S202 determines the dynamic allocation bitrate of video frames based on the basic coding strategy and the overall base bitrate of the video frames.
[0048] In this embodiment, dynamic bitrate allocation refers to the redundant bitrate that can be allocated to video frames, which is also the redundancy of the overall bitrate budget. The overall base bitrate refers to the base bitrate that can be allocated to a video frame.
[0049] An alternative approach is to obtain the dynamically allocated bitrate of video frames based on a bitrate allocation model, according to the basic coding strategy and the overall base bitrate of the video frames. The bitrate allocation model is obtained by pre-training a neural network based on the bitrates corresponding to the encoded video.
[0050] Another option is to use a basic coding strategy to calculate the overall bitrate ceiling of the video frame; and then determine the dynamic bitrate allocation of the video frame based on the overall bitrate ceiling and the overall basic bitrate.
[0051] The overall bitrate cap refers to the upper limit of the overall bitrate budget corresponding to the video frame.
[0052] Specifically, the overall bitrate upper limit of the video frame can be calculated according to the basic coding strategy, and the overall base bitrate of the video frame can be obtained based on the complexity of the video frame, such as intra-frame prediction error and motion vector amplitude. Then, the difference between the overall bitrate upper limit and the overall base bitrate is used as the dynamic bitrate allocation of the video frame.
[0053] Understandably, the first step is to determine the dynamic allocation bitrate of video frames, that is, to determine the redundancy of the overall bitrate budget for video encoding, in order to facilitate subsequent video encoding.
[0054] S203 determines the additional bitrate of the target region based on the overall base bitrate of the video frame, as well as the area ratio of the target region, visual sensitivity, and texture complexity.
[0055] The additional bitrate refers to the extra bitrate required for the target region. The region area percentage refers to the percentage of the target region within a video frame.
[0056] An alternative approach is to use an additional bitrate determination model to determine the additional bitrate of a target region based on the overall base bitrate of the video frames, the area proportion of the target region, visual sensitivity, and texture complexity. This additional bitrate determination model can be trained on a neural network based on the additional bitrate of the already encoded video.
[0057] Another option is to calculate the reference bitrate of the target region based on the overall base bitrate of the video frame and the area ratio of the target region; and determine the additional bitrate of the target region based on the visual sensitivity, texture complexity, and reference bitrate of the target region.
[0058] The reference bitrate refers to the bitrate when encoding based on conventional quantization parameters.
[0059] Specifically, the reference bitrate of the target region is obtained by multiplying the overall base bitrate of the video frame by the area ratio of the target region. Then, the visual sensitivity of the target region is multiplied by the first weighting coefficient to obtain the visual sensitivity term. The result of subtracting the texture complexity from 1 is multiplied by the second weighting coefficient to obtain the texture complexity term. The result of adding the visual sensitivity term and the texture complexity term is multiplied by the reference bitrate to obtain the additional bitrate of the target region.
[0060] Understandably, by determining the additional bitrate for the target region within the overall bitrate budget constraint of video encoding, the foundation is laid for subsequent bitrate allocation for the target region.
[0061] S204: Based on the dynamically allocated bitrate and the additional bitrate, the bitrate is allocated to the target area to obtain the bitrate allocation result.
[0062] An alternative approach is to allocate a corresponding bitrate to the target region from the dynamically allocated bitrate based on the additional bitrate of the target region, thereby obtaining the bitrate classification result.
[0063] Another alternative approach is to allocate bitrate to the target region based on the dynamically allocated bitrate and the additional bitrate, and obtain the bitrate allocation result. This includes: prioritizing the visual sensitivity of the target region, allocating the dynamically allocated bitrate of the video frame to the target region based on the additional bitrate of the target region, and obtaining the bitrate allocation result.
[0064] Specifically, priority is given to target areas with high visual sensitivity. The dynamic allocation bitrate of video frames is allocated to the target areas based on the additional bitrate of the target areas to obtain the bitrate allocation result.
[0065] It is understandable that the visual sensitivity of the target area is prioritized, and the limited bitrate is tilted towards the target area, while ensuring the basic encoding quality of non-target areas.
[0066] For example, prioritizing the visual sensitivity of the target region, the dynamic allocation bitrate of video frames is allocated to the target region based on the additional bitrate of the target region. This includes: if the dynamic allocation bitrate is greater than or equal to the additional bitrate corresponding to the target region with the highest visual sensitivity, then the additional bitrate is allocated to that target region, and the difference between the dynamic allocation bitrate and the additional bitrate is used as the new dynamic allocation bitrate; from the remaining target regions, the region with the highest visual sensitivity is selected as the new target region, and allocation continues until the dynamic allocation bitrate is exhausted.
[0067] If the dynamically allocated bitrate is less than the additional bitrate corresponding to the target area with the highest visual sensitivity, the target areas are scored and sorted according to their visual sensitivity, and the dynamically allocated bitrate is given priority to the target areas with high scores, until the dynamically allocated bitrate is used up, ensuring that the overall bitrate does not exceed the overall bitrate limit.
[0068] It is understandable that the bitrate of the target area is balanced to facilitate subsequent video encoding and improve video quality.
[0069] S205, based on the bitrate allocation result, adjusts the quantization parameters of the target region and non-target region in the video frame to obtain differentiated coding parameters.
[0070] S206 uses a video encoder to encode video frames according to differentiated encoding parameters.
[0071] The technical solution provided in this disclosure identifies target regions by performing sensitive region identification on video frames. Based on the basic coding strategy and the overall base bitrate of the video frame, a dynamically allocated bitrate is determined. An additional bitrate for the target region is determined based on the overall base bitrate, the area ratio of the target region, its visual sensitivity, and texture complexity. Bitrate allocation is performed on the target region based on the dynamically allocated bitrate and the additional bitrate, resulting in a bitrate allocation result. Based on the bitrate allocation result, quantization parameters are adjusted for both target and non-target regions in the video frame to obtain differentiated coding parameters. A video encoder is then used to encode the video frame according to these differentiated coding parameters. This technical solution employs a dynamic bitrate allocation strategy, which, under the constraint of the overall bitrate budget, tilts the bitrate towards highly sensitive, pixelation-prone areas, avoiding the bitrate waste and insufficient bitrate in sensitive areas caused by globally uniform allocation, thereby improving video quality.
[0072] The technical solution disclosed herein overcomes the shortcomings of existing low bitrate video encoding technologies, such as lack of targeted bitrate allocation, insufficient accuracy in sensitive area identification, and single encoding optimization strategy, which lead to severe distortion in subjectively prone mosaic areas. By accurately identifying specific subjectively sensitive areas such as smoke, sky, skin, and light that are sensitive to the human eye and prone to mosaic, and by adopting dynamic bitrate allocation and adaptive encoding optimization strategies, the mosaic distortion in sensitive areas is effectively reduced without significantly increasing the overall bitrate overhead, thereby improving the subjective viewing experience of low bitrate videos.
[0073] Figure 3 is a schematic diagram of a video frame encoding device according to an embodiment of this disclosure; this embodiment is applicable to situations where mosaic distortion can be avoided in low bitrate extreme compression scenarios. The device can be implemented in software and / or hardware and can be integrated into electronic devices carrying video frame encoding functions, such as servers. As shown in Figure 3, the video frame encoding device 300 provided in this embodiment includes: a target region determination module 301, used to identify sensitive regions of the video frame to obtain target regions; a bitrate allocation module 302, used to allocate bitrate to the target region based on the basic encoding strategy, the overall basic bitrate of the video frame, and the area ratio, visual sensitivity, and texture complexity of the target region, to obtain the bitrate allocation result of the target region; a parameter adjustment module 303, used to adjust the quantization parameters of the target region and non-target regions in the video frame based on the bitrate allocation result, to obtain differentiated encoding parameters; and a video encoding module 304, used to encode the video frame using a video encoder according to the differentiated encoding parameters.
[0074] The technical solution provided in this disclosure identifies target regions by performing sensitive region identification on video frames. Based on a basic encoding strategy, the overall base bitrate of the video frame, and the area ratio, visual sensitivity, and texture complexity of the target region, bitrate allocation is performed on the target region to obtain a bitrate allocation result. Based on the bitrate allocation result, quantization parameters are adjusted for target and non-target regions in the video frame to obtain differentiated encoding parameters. A video encoder is then used to encode the video frame according to these differentiated encoding parameters. This technical solution, by combining the overall base bitrate of the video frame with the area ratio, visual sensitivity, and texture complexity of the target region for bitrate allocation, and then performing differentiated encoding on target and non-target regions in the video frame, can improve the image quality of visually sensitive mosaic areas in video frames.
[0075] Furthermore, the target region determination module 301 is used to: perform semantic segmentation on video frames based on preset sensitive categories to obtain initial regions belonging to sensitive categories; calculate the texture complexity of the initial regions and filter candidate regions that are prone to mosaic based on the texture complexity; determine the visual sensitivity of the candidate regions and filter target regions based on the visual sensitivity.
[0076] Furthermore, the bitrate allocation module 302 includes: a dynamic bitrate allocation determination unit, used to determine the dynamic bitrate allocation of the video frame based on the basic coding strategy and the overall base bitrate of the video frame; an additional bitrate determination unit, used to determine the additional bitrate of the target region based on the overall base bitrate of the video frame, as well as the area ratio, visual sensitivity, and texture complexity of the target region; and a bitrate allocation result determination unit, used to allocate bitrate to the target region based on the dynamic bitrate allocation and the additional bitrate, to obtain the bitrate allocation result.
[0077] Furthermore, based on the basic coding strategy, the dynamic bitrate allocation determination unit is used to: calculate the overall bitrate upper limit of the video frame using the basic coding strategy; and determine the dynamic bitrate allocation of the video frame based on the overall bitrate upper limit and the overall basic bitrate.
[0078] Furthermore, the additional bitrate determination unit is used to: calculate the reference bitrate of the target region based on the overall base bitrate of the video frame and the area ratio of the target region; and determine the additional bitrate of the target region based on the visual sensitivity, texture complexity, and reference bitrate of the target region.
[0079] Furthermore, the bitrate allocation result determination unit is used to: prioritize the visual sensitivity of the target area, allocate the dynamic allocation bitrate of the video frame to the target area according to the additional bitrate of the target area, so as to obtain the bitrate allocation result.
[0080] Furthermore, the bitrate allocation result determination unit is used to: if the dynamically allocated bitrate is greater than or equal to the additional bitrate corresponding to the target region with the highest visual sensitivity, then allocate an additional bitrate to the target region and use the difference between the dynamically allocated bitrate and the additional bitrate as the new dynamically allocated bitrate; select the region with the highest visual sensitivity from the remaining target regions as the new target region, and continue to allocate bitrate until the dynamically allocated bitrate is exhausted.
[0081] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0082] Figure 4 is a block diagram of an electronic device for implementing the video frame encoding method of an embodiment of the present disclosure. Figure 4 illustrates a schematic block diagram of an example electronic device 400 that can be used to implement embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0083] As shown in Figure 4, the electronic device 400 includes a computing unit 401, which can perform various appropriate actions and processes based on a computer program stored in a read-only memory (ROM) 402 or a computer program loaded into a random access memory (RAM) 403 from a storage unit 408. The RAM 403 may also store various programs and data required for the operation of the electronic device 400. The computing unit 401, ROM 402, and RAM 403 are interconnected via a bus 404. An input / output (I / O) interface 405 is also connected to the bus 404.
[0084] Multiple components in electronic device 400 are connected to I / O interface 405, including: input unit 406, such as keyboard, mouse, etc.; output unit 407, such as various types of displays, speakers, etc.; storage unit 408, such as disk, optical disk, etc.; and communication unit 409, such as network card, modem, wireless transceiver, etc. Communication unit 409 allows electronic device 400 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0085] The computing unit 401 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 401 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 401 performs the various methods and processes described above, such as video frame encoding methods. For example, in some embodiments, the video frame encoding method may be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 408. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 400 via ROM 402 and / or communication unit 409. When the computer program is loaded into RAM 403 and executed by the computing unit 401, one or more steps of the video frame encoding method described above may be performed. Alternatively, in other embodiments, the computing unit 401 may be configured to perform the video frame encoding method by any other suitable means (e.g., by means of firmware).
[0086] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations may include: implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0087] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0088] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0089] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0090] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0091] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.
[0092] Artificial intelligence (AI) is the study of enabling computers to simulate certain human thought processes and intelligent behaviors (such as learning, reasoning, thinking, and planning). It encompasses both hardware and software technologies. AI hardware technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, and big data processing. AI software technologies mainly include computer vision, speech recognition, natural language processing, machine learning / deep learning, big data processing, and knowledge graph technologies.
[0093] Cloud computing refers to a technology system that enables access to a shared pool of physical or virtual resources via a network. These resources can include servers, operating systems, networks, software, applications, and storage devices, and can be deployed and managed on demand and in a self-service manner. Cloud computing technology can provide efficient and powerful data processing capabilities for applications such as artificial intelligence and blockchain, as well as for model training.
[0094] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.
[0095] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.
Claims
1. A video frame encoding method, comprising: Sensitive region identification is performed on video frames to obtain the target region; Based on the basic coding strategy, the overall base bitrate of the video frame, and the area ratio, visual sensitivity, and texture complexity of the target region, bitrate allocation is performed for the target region to obtain the bitrate allocation result of the target region. Based on the bitrate allocation result, the quantization parameters of the target region and non-target region in the video frame are adjusted to obtain differentiated coding parameters; the video frame is then encoded using a video encoder according to the differentiated coding parameters.
2. The method according to claim 1, wherein, The step of identifying sensitive regions in a video frame to obtain a target region includes: performing semantic segmentation on the video frame based on a preset sensitive category to obtain an initial region belonging to the sensitive category; calculating the texture complexity of the initial region and filtering candidate regions prone to mosaic effects from the initial region based on the texture complexity; determining the visual sensitivity of the candidate regions and filtering target regions from the candidate regions based on the visual sensitivity.
3. The method according to claim 1, wherein, The step of allocating bitrate to the target region based on the basic coding strategy, the overall base bitrate of the video frame, and the area ratio, visual sensitivity, and texture complexity of the target region to obtain the bitrate allocation result for the target region includes: determining the dynamic allocation bitrate of the video frame based on the basic coding strategy and the overall base bitrate of the video frame; determining the additional bitrate of the target region based on the overall base bitrate of the video frame, and the area ratio, visual sensitivity, and texture complexity of the target region; and allocating bitrate to the target region based on the dynamic allocation bitrate and the additional bitrate to obtain the bitrate allocation result.
4. The method according to claim 3, wherein, The step of determining the dynamic allocation bitrate of the video frame based on the basic coding strategy and the overall basic bitrate of the video frame includes: using the basic coding strategy to calculate the overall bitrate upper limit of the video frame; and determining the dynamic allocation bitrate of the video frame based on the overall bitrate upper limit and the overall basic bitrate.
5. The method according to claim 3, wherein, The step of determining the additional bitrate of the target region based on the overall base bitrate of the video frame, the area ratio of the target region, visual sensitivity, and texture complexity includes: calculating the reference bitrate of the target region based on the overall base bitrate of the video frame and the area ratio of the target region; and determining the additional bitrate of the target region based on the visual sensitivity, texture complexity, and reference bitrate of the target region.
6. The method according to claim 3, wherein, Based on the dynamically allocated bitrate and the additional bitrate, bitrate allocation is performed for the target region to obtain a bitrate allocation result, including: prioritizing the visual sensitivity of the target region, allocating the dynamically allocated bitrate of the video frame to the target region according to the additional bitrate of the target region, so as to obtain a bitrate allocation result.
7. The method according to claim 6, wherein, Prioritizing the visual sensitivity of the target region, the dynamic allocation bitrate of video frames is allocated to the target region based on the additional bitrate of the target region. This includes: if the dynamic allocation bitrate is greater than or equal to the additional bitrate corresponding to the target region with the highest visual sensitivity, then the additional bitrate is allocated to that target region, and the difference between the dynamic allocation bitrate and the additional bitrate is used as the new dynamic allocation bitrate; from the remaining target regions, the region with the highest visual sensitivity is selected as the new target region, and allocation continues until the dynamic allocation bitrate is exhausted.
8. A video frame encoding apparatus, comprising: The target region determination module is used to identify sensitive regions in video frames to obtain the target region; The bitrate allocation module is used to allocate bitrate to the target region based on the basic encoding strategy, the overall basic bitrate of the video frame, the area ratio of the target region, visual sensitivity, and texture complexity, and obtain the bitrate allocation result of the target region. The parameter adjustment module is used to adjust the quantization parameters of the target region and non-target region in the video frame based on the bitrate allocation result, so as to obtain differentiated coding parameters; The video encoding module is used to encode the video frames according to the differentiated encoding parameters using a video encoder.
9. An electronic device, comprising: At least one processor; The at least one processor is also connected in communication with a memory, wherein the memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the video frame encoding method of any one of claims 1-7.
10. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to execute the video frame encoding method according to any one of claims 1-7.
11. A computer program product comprising a computer program that, when executed by a processor, implements the video frame encoding method according to any one of claims 1-7.