Video coding method, device, equipment, medium and product
By classifying the brightness features and optimizing the quantization parameters of each frame of HDR video data, the problem of neglecting visual perception characteristics in traditional video coding standards is solved, achieving efficient utilization of bit resources and improved video quality.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-18
- Publication Date
- 2026-04-03
AI Technical Summary
Traditional video coding standards neglect visual perception characteristics when encoding HDR videos, resulting in low bit resource utilization efficiency and degraded video coding quality.
By acquiring the luminance feature information of each frame, the frame data is classified based on the luminance feature information, and the corresponding encoding strategy and quantization parameters are determined according to the frame category for video encoding.
It improves the efficiency of bit resource utilization while ensuring video encoding quality and avoiding visual discontinuity caused by differences in frame-level parameters.
Smart Images

Figure CN121792736A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of video coding technology, and in particular to a video coding method, apparatus, device, medium and product. Background Technology
[0002] Currently, when encoding HDR (High Dynamic Range) video, traditional video coding standards such as H.265 / HEVC (High Efficiency Video Coding) are typically used directly. Traditional video coding standards, when optimizing rate-distortion, primarily rely on the mathematical characteristics of the signal (such as transform coefficient energy, motion vector size, and texture complexity), severely neglecting visual perception characteristics. This leads to situations where, in scenes where visual sensitivity decreases, the encoder allocates a large number of bits but fails to effectively improve subjective quality, resulting in low bit resource utilization efficiency.
[0003] Therefore, how to improve the efficiency of bit resource utilization while ensuring video encoding quality is a technical problem that needs to be solved by those skilled in the art. Summary of the Invention
[0004] In view of this, the purpose of this application is to provide a video encoding method, apparatus, device, medium, and product that can improve bit resource utilization efficiency while ensuring video encoding quality. The specific solution is as follows:
[0005] In a first aspect, this application provides a video coding method, including:
[0006] Obtain the luminance feature information corresponding to each frame of data in the video data to be encoded;
[0007] The frame category of each frame of data is determined based on the brightness feature information;
[0008] The encoding strategy corresponding to the frame category is determined based on a preset correspondence.
[0009] The quantization parameters corresponding to each frame of data are determined according to the encoding strategy described above.
[0010] Each frame of data is encoded based on the quantization parameters corresponding to each frame of data to obtain video encoded data.
[0011] Optionally, obtain the luminance feature information corresponding to each frame of data in the video data to be encoded, including:
[0012] Obtain the maximum and average brightness of each frame in the video data to be encoded.
[0013] Optionally, determining the frame category of each frame of data based on the brightness feature information includes:
[0014] The frame category of each frame of data is determined based on the maximum brightness, the average brightness, the preset maximum brightness classification threshold, and the preset average brightness classification threshold.
[0015] Optionally, the preset average brightness classification threshold includes a first threshold and a second threshold, wherein the first threshold is less than the second threshold. Based on the maximum brightness, the average brightness, and the preset maximum brightness classification threshold and the preset average brightness classification threshold, the frame category of each frame of data is determined, including:
[0016] For any frame of data, determine whether the average brightness of the frame of data is less than the first threshold. If it is less than the first threshold, determine that the frame of data is a dark frame.
[0017] If the average brightness of the frame data is greater than or equal to the first threshold, then it is determined whether the average brightness of the frame data is greater than the second threshold and whether the maximum brightness of the frame data is greater than the preset maximum brightness division threshold. If the average brightness of the frame data is greater than the second threshold and the maximum brightness of the frame data is greater than the preset maximum brightness division threshold, then the frame category of the frame data is determined to be a bright frame.
[0018] If the average brightness of the frame data is less than or equal to the second threshold, and the maximum brightness of the frame data is greater than the preset maximum brightness classification threshold, then the frame category of the frame data is determined to be a mid-gray detail frame.
[0019] If the average brightness of the frame data is greater than or equal to the second threshold, and the maximum brightness of the frame data is less than or equal to the preset maximum brightness division threshold, then the frame category of the frame data is determined to be a smooth frame.
[0020] Optionally, the quantization parameters corresponding to each frame of data are determined according to the encoding strategy, including:
[0021] The quantization parameter offset for each frame of data is determined according to the encoding strategy described above.
[0022] The quantization parameters for each frame of data are determined based on the quantization parameter offset and the standard quantization parameters.
[0023] Optionally, after determining the quantization parameter offset corresponding to each frame of data according to the encoding strategy, the method further includes:
[0024] The quantization parameter offsets corresponding to each frame of data are used to form an offset sequence;
[0025] The offset sequence is smoothed and filtered to obtain the filtered offset for each frame of data;
[0026] Accordingly, the quantization parameters for each frame of data are determined based on the quantization parameter offset and the standard quantization parameters, including:
[0027] The quantization parameters for each frame of data are determined based on the filtered offset and the standard quantization parameters.
[0028] Secondly, this application provides a video encoding apparatus, comprising:
[0029] The luminance feature acquisition module is used to acquire the luminance feature information corresponding to each frame of data in the video data to be encoded;
[0030] A frame category determination module is used to determine the frame category of each frame of data based on the brightness feature information.
[0031] The encoding strategy determination module is used to determine the encoding strategy corresponding to the frame category based on a preset correspondence.
[0032] The quantization parameter determination module is used to determine the quantization parameters corresponding to each frame of data according to the encoding strategy.
[0033] The video data encoding module is used to encode each frame of data based on the quantization parameters corresponding to each frame of data to obtain video encoded data.
[0034] Thirdly, this application provides an electronic device, including a memory and a processor, wherein:
[0035] The memory is used to store computer programs;
[0036] The processor is used to execute the computer program to implement the aforementioned video encoding method.
[0037] Fourthly, this application provides a computer-readable storage medium for storing a computer program, wherein the computer program, when executed by a processor, implements the aforementioned video encoding method.
[0038] Fifthly, this application provides a computer program product, including a computer program / instructions that, when executed by a processor, implement the aforementioned video encoding method.
[0039] As can be seen from the above scheme, the present invention provides a video encoding method, including: acquiring luminance feature information corresponding to each frame of data in the video data to be encoded; determining the frame category of each frame of data based on the luminance feature information; determining the encoding strategy corresponding to the frame category based on a preset correspondence; determining the quantization parameter corresponding to each frame of data according to the encoding strategy; and encoding each frame of data based on the quantization parameter corresponding to each frame of data to obtain video encoded data.
[0040] As can be seen, the beneficial effects of this application are as follows: Each frame of video data is classified based on its luminance feature information; then, the encoding strategy corresponding to the frame category is determined based on a preset correspondence; subsequently, the quantization parameters corresponding to each frame are determined according to the encoding strategy; and finally, each frame is encoded based on its quantization parameters to obtain the encoded video data. In this way, by determining the corresponding encoding strategy for each frame based on its luminance features, and then determining the quantization parameters used for encoding each frame, the visual perception features of each frame are fully considered, and encoding is performed separately. This approach can improve bit resource utilization efficiency while ensuring video encoding quality.
[0041] Correspondingly, the video encoding apparatus, device, and readable storage medium provided in this application also have the above-mentioned technical effects. Attached Figure Description
[0042] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of this application. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.
[0043] Figure 1 A flowchart of a video encoding method provided in this application embodiment;
[0044] Figure 2 This is a schematic diagram of a video encoding device provided in an embodiment of this application;
[0045] Figure 3 This is a structural diagram of an electronic device provided in an embodiment of this application. Detailed Implementation
[0046] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0047] Currently, HDR video encoding typically employs traditional video coding standards such as H.265 / HEVC. These standards, when optimizing rate-distortion, primarily rely on the mathematical characteristics of the signal (such as transform coefficient energy, motion vector size, and texture complexity), severely neglecting the visual perception characteristics of HDR content. For example, in HDR content, certain bright scenes (such as water surfaces directly exposed to sunlight) exhibit large local brightness gradients, leading to a significant increase in residual energy and thus requiring more bit resources. However, based on visual contrast sensitivity characteristics, the human eye's sensitivity to subtle texture differences decreases against extremely bright backgrounds. This means that in these scenes, the large number of bits allocated by the encoder cannot effectively improve subjective quality, resulting in low bit resource utilization efficiency. Therefore, this application provides a video coding scheme that can improve bit resource utilization efficiency while ensuring video encoding quality.
[0048] See Figure 1 As shown in the figure, this application discloses a video encoding method, including:
[0049] Step S11: Obtain the luminance feature information corresponding to each frame of data in the video data to be encoded.
[0050] In an optional implementation, embodiments of this application can obtain the luminance feature information corresponding to each frame of data from the dynamic metadata corresponding to the video data to be encoded. That is, if the dynamic metadata corresponding to the video data to be encoded contains luminance feature information corresponding to each frame of data, then the luminance feature information corresponding to each frame of data is obtained from the dynamic metadata. For example, the dynamic metadata of HDR VIVID (i.e., Vivid Image Representation) contains display luminance information for each frame of data, and the display luminance information corresponding to each frame of data, i.e., the luminance feature information, can be obtained from the dynamic metadata. The dynamic metadata contained in HDR VIVID contains display luminance information for each frame, which is typically only used for tone mapping at the display end. This application integrates this information into the encoding decision, making full use of it.
[0051] This involves obtaining the luminance feature information corresponding to each frame of the video data to be encoded, including: obtaining the maximum luminance and average luminance corresponding to each frame of the video data to be encoded. The maximum luminance is the highest luminance among all pixels in the frame, and the average luminance is the average luminance of all pixels in the frame.
[0052] In other words, brightness feature information can include maximum brightness and average brightness. Dynamic metadata contains the maximum brightness and average brightness of each frame of data. The maximum brightness and average brightness corresponding to each frame of data can be obtained from the dynamic metadata to obtain brightness feature information.
[0053] Step S12: Determine the frame category of each frame of data based on the brightness feature information.
[0054] That is, in this embodiment of the application, each frame of data is classified according to its brightness feature information. Classification can be based on brightness feature information and a preset threshold. In an optional implementation, the frame category of each frame of data can be determined based on the maximum brightness, the average brightness, and preset maximum brightness and average brightness classification thresholds.
[0055] The preset average brightness classification threshold includes a first threshold and a second threshold, where the first threshold is less than the second threshold. Based on the maximum brightness, the average brightness, and the preset maximum brightness classification threshold and the preset average brightness classification threshold, the frame category of each frame of data is determined, including:
[0056] Furthermore, for any given frame of data, it is determined whether the average brightness of the frame is less than the first threshold. If it is less than the first threshold, the frame type of the frame is determined to be a dark frame. If it is greater than or equal to the first threshold, it is determined whether the average brightness of the frame is greater than the second threshold and whether the maximum brightness of the frame is greater than the preset maximum brightness threshold. If the average brightness of the frame is greater than the second threshold and the maximum brightness of the frame is greater than the preset maximum brightness threshold, the frame type of the frame is determined to be a bright frame. If the average brightness of the frame is less than or equal to the second threshold and the maximum brightness of the frame is greater than the preset maximum brightness threshold, the frame type of the frame is determined to be a mid-gray detail frame. If the average brightness of the frame is greater than or equal to the second threshold and the maximum brightness of the frame is less than or equal to the preset maximum brightness threshold, the frame type of the frame is determined to be a smooth frame.
[0057] Step S13: Determine the encoding strategy corresponding to the frame category based on the preset correspondence.
[0058] The embodiments of this application can preset encoding strategies corresponding to different frame categories, that is, different frame categories correspond to different encoding strategies, and the encoding strategy may include a quantization parameter offset determination strategy.
[0059] Step S14: Determine the quantization parameters corresponding to each frame of data according to the encoding strategy.
[0060] In this embodiment of the application, the quantization parameter offset corresponding to each frame of data can be determined according to the encoding strategy; the quantization parameter corresponding to each frame of data can be determined based on the quantization parameter offset and the standard quantization parameter.
[0061] In this system, bright frames correspond to a first quantization parameter offset, which is a positive number; dark frames correspond to a second quantization parameter offset, which is a negative number; mid-gray detail frames can have a quantization parameter offset of 0; and smooth frames correspond to a third quantization parameter offset, which is a positive number and less than the first quantization parameter offset. The standard quantization parameter can be set by the user or calculated by the encoder given a bitrate. The sum of the quantization parameter offset and the standard quantization parameter is the quantization parameter corresponding to that frame. For example, if the standard quantization parameter is 32, the third quantization parameter offset can be 1 or 2, corresponding to quantization parameters 33 or 34; if the first quantization parameter offset is 3, the corresponding quantization parameter is 35; and if the second quantization parameter offset is -1 or -2, the corresponding quantization parameter is 30 or 31. The aforementioned quantization parameter offset and standard quantization parameter refer to the quantization parameter offset and standard quantization parameter corresponding to brightness. In an optional implementation, for bright frames, the QP offset of the chroma component can also be reduced to protect color from distortion. For dark frames, the strength of the deblocking filter can be increased to smooth out the blocking effect that may be caused by quantization in dark areas.
[0062] In an optional implementation, after determining the quantization parameter offset corresponding to each frame of data according to the encoding strategy, the method may further include: constructing an offset sequence from the quantization parameter offsets corresponding to each frame of data; performing smoothing filtering on the offset sequence to obtain the filtered offset corresponding to each frame of data; and correspondingly, determining the quantization parameter corresponding to each frame of data based on the quantization parameter offset and the standard quantization parameter, including: determining the quantization parameter corresponding to each frame of data based on the filtered offset and the standard quantization parameter.
[0063] The offset sequence can be smoothed using a temporal low-pass filter. For example, a moving average filter or a Gaussian filter can be used, with a fixed window size centered on the current frame, to smooth the offset sequence. This ensures that parameter changes between adjacent frames are continuous in transition regions (such as scene switching points), effectively eliminating abrupt changes in image quality and guaranteeing a smooth visual experience without significantly deviating from the optimal encoding strategy for a single frame.
[0064] Step S15: Encode each frame of data based on the quantization parameters corresponding to each frame of data to obtain video encoded data.
[0065] That is, in the embodiments of this application, each frame of data is encoded according to the quantization parameters corresponding to each frame to obtain video encoded data.
[0066] As can be seen, the embodiments of this application classify each frame of data based on the luminance feature information corresponding to each frame in the video data to be encoded. Then, based on a preset correspondence, the encoding strategy corresponding to the frame category is determined. Subsequently, the quantization parameters corresponding to each frame of data are determined according to the encoding strategy, and then each frame of data is encoded based on the quantization parameters corresponding to each frame of data to obtain the video encoded data. In this way, by determining the corresponding encoding strategy for each frame of data according to the luminance features of each frame of data, and then determining the quantization parameters used for each frame of data during encoding, the visual perception features of each frame of data are fully considered, and the data is encoded separately. This can improve the efficiency of bit resource utilization while ensuring the quality of video encoding.
[0067] Furthermore, in this embodiment, before encoding, the luminance metadata of each frame can be obtained through a pre-analysis channel. Then, based on this metadata, each frame is divided into multiple predefined perceptual categories, i.e., frame categories. Next, a customized encoding strategy is selected for each frame based on the classification results. Simultaneously, to address the visual discontinuity problem that may be caused by differences in frame-level parameters, a parameter smoothing transition mechanism is introduced. Finally, the encoder encodes based on the smoothed parameters. In this way, pre-analysis is performed before video encoding, pre-acquiring the dynamic metadata of each frame. Using the maximum and average brightness of the displayed content, the visual importance of each frame is classified, and based on this classification, a differentiated set of encoding parameters is adaptively configured for different frames. Ultimately, this significantly improves the compression efficiency of HDR video while maintaining visual coherence. By deeply integrating the visual perception characteristics of HDR with the video encoding process, and by using pre-known frame-level luminance metadata to classify the visual importance of each frame, the limited bitrate is shifted from areas insensitive to the human eye to highly sensitive areas, maximizing global perceptual quality.
[0068] Among them, in this application, a perception classifier is constructed, defining a two-dimensional classification space. Dimension one is based on the maximum display brightness Lmax in the dynamic metadata; dimension two is based on the average display brightness Lavg in the dynamic metadata. Combining these two dimensions, the frames to be encoded can be divided into the following typical frame categories: a. High dynamic range bright frames (i.e., the aforementioned bright frames), characteristics: both the maximum display brightness and the average display brightness are high. Scene examples: beaches, snowfields, sky shots under sunlight. Visual characteristics: The entire picture is very bright. The human eye is not sensitive to high-light details but is sensitive to color saturation, and the overall visual sensitivity is low. b. Dark field frames, characteristics: The average display brightness is very low. Scene examples: night street scenes, dim interiors. Visual characteristics: The average brightness of the entire picture is very low. The human eye is extremely sensitive to noise and blocking effects in dark areas. The visual sensitivity is high. c. Medium gray detail frames, characteristics: frames that do not belong to a and b, and the maximum display brightness is high. Scene examples: outdoors on cloudy days, interiors with uniform lighting. Visual characteristics: The picture brightness is moderate, but the dynamic range is relatively wide, containing rich details from dark to bright areas. The human eye has high visual sensitivity and is sensitive to the loss of details and textures. d. Smooth frames, characteristics: frames that do not belong to a and b, and the maximum display brightness is low. Scene examples: clean walls, foggy days, soft-focus backgrounds, solid color panels. Visual characteristics: Include two situations where the picture brightness is moderate but the dynamic range is very narrow, and the overall is very bright but without glaring highlights. The content is smooth, lacking strong contrast and fine textures.
[0069] Furthermore, pre-analysis and frame classification are performed. For the input sequence, according to the metadata tuple {Lavg, Lmax} corresponding to each frame, using the perception classifier, and based on the preset thresholds T1_low (i.e., the aforementioned first threshold), T1_high (i.e., the aforementioned second threshold), and T2 (i.e., the aforementioned preset maximum brightness division threshold), each frame is mapped to a certain category of the perception classifier. The classification logic is organized as follows: Judge Lavg: If Lavg < T1_low, then judge it as a dark field frame, otherwise, continue to judge. Judge Lavg and Lmax: If Lavg > T1_high and Lmax > T2, then judge it as a high dynamic range bright frame, otherwise, continue to judge. Judge Lmax: If Lmax > T2, then judge it as a medium gray detail frame, otherwise, judge it as a smooth frame.
[0070] Furthermore, the encoding strategy is determined, with a customized encoding strategy preset for each category. The customized encoding strategies are as follows: a. High dynamic range bright frames: A positive QP (Quantization Parameter) offset value is set for the entire frame (i.e., coarser quantization), while the QP offset of the chroma component is relatively reduced to protect color from distortion. Video data YCBCR contains luminance and chroma components. This application mainly adjusts the luminance QP, but chroma also has a QP, which can be adjusted selectively in bright frames. b. Dark frames: A negative QP offset value is set for the entire frame (i.e., finer quantization), while the strength of the deblocking filter can be appropriately enhanced to smooth out the block effect that may be caused by quantization in dark areas. c. Mid-gray detail frames: Standard QP is used without large offset, i.e., the offset is 0. d. Smooth frames: A positive QP offset is used, which is smaller than the offset corresponding to luminance in a, resulting in moderate bitrate savings.
[0071] Furthermore, encoding smoothing is performed. Based on the pre-analysis and classification results, a QP offset value corresponding to brightness can be determined for each frame. This value reflects the independent optimal strategy for each frame. To avoid inter-frame quality fluctuations and visual discontinuities caused by applying different encoding strategies between frames, a parameter smoothing transition mechanism is introduced to filter the original sequence QP offset values generated by classification. Specifically, after the encoding strategy is determined, this step-like parameter sequence is not used immediately. Instead, a temporal low-pass filter (such as a moving average filter or a Gaussian filter with a fixed window size centered on the current frame) is applied. This filtering process takes the original QP offset value generated by the encoding strategy as input and outputs a smooth parameter curve, thereby ensuring that the parameter changes between adjacent frames are continuous in transition regions (such as scene switching points). This effectively eliminates the abrupt changes in image quality without significantly deviating from the optimal encoding strategy for a single frame, ensuring a smooth visual experience.
[0072] Finally, the video is encoded using the smoothed parameters.
[0073] In this way, by classifying frames based on known metadata, the encoder no longer mechanically processes each frame but instead encodes based on the visual intent of the current frame and adopts the most appropriate strategy. By directly applying different QP offsets to different categories of frames, HDR visual characteristics are transformed into concise and efficient encoding instructions, significantly improving efficiency. A simple low-pass filter smooths the classified QP curves, perfectly resolving the flickering problem caused by frame-level quality differences with extremely low computational cost, ensuring visual continuity. The implementation is simple, requiring no modification to the complex logic at the encoding block level; only different parameter sets need to be passed to the frame-level interface, making it easy to implement on existing encoders, while the resulting bitrate savings (or quality improvements at the same bitrate) are very significant.
[0074] Further, see Figure 2 As shown, this application embodiment provides a video encoding apparatus, including:
[0075] The brightness feature acquisition module 11 is used to acquire the brightness feature information corresponding to each frame of data in the video data to be encoded;
[0076] Frame category determination module 12 is used to determine the frame category of each frame of data based on the brightness feature information;
[0077] Encoding strategy determination module 13 is used to determine the encoding strategy corresponding to the frame category based on a preset correspondence relationship;
[0078] The quantization parameter determination module 14 is used to determine the quantization parameters corresponding to each frame of data according to the encoding strategy.
[0079] The video data encoding module 15 is used to encode each frame of data based on the quantization parameters corresponding to each frame of data to obtain video encoded data.
[0080] Among them, the brightness feature acquisition module 11 can be specifically used to acquire the maximum brightness and average brightness corresponding to each frame of data in the video data to be encoded.
[0081] The frame category determination module 12 can be specifically used to determine the frame category of each frame of data based on the maximum brightness, the average brightness, and the preset maximum brightness division threshold and the preset average brightness division threshold.
[0082] The preset average brightness classification threshold includes a first threshold and a second threshold. The frame category determination module 12 can be specifically used to: for any frame data, determine whether the average brightness of the frame data is less than the first threshold; if it is less than the first threshold, determine that the frame category of the frame data is a dark frame; if it is greater than or equal to the first threshold, determine whether the average brightness of the frame data is greater than the second threshold and whether the maximum brightness of the frame data is greater than the preset maximum brightness classification threshold; if the average brightness of the frame data is greater than the second threshold and the maximum brightness of the frame data is greater than the preset maximum brightness classification threshold, determine that the frame category of the frame data is a bright frame; if the average brightness of the frame data is less than or equal to the second threshold and the maximum brightness of the frame data is greater than the preset maximum brightness classification threshold, determine that the frame category of the frame data is a mid-gray detail frame; if the average brightness of the frame data is greater than or equal to the second threshold and the maximum brightness of the frame data is less than or equal to the preset maximum brightness classification threshold, determine that the frame category of the frame data is a smooth frame.
[0083] Quantization parameter determination module 14 may specifically include:
[0084] The offset determination submodule is used to determine the quantization parameter offset corresponding to each frame of data according to the encoding strategy.
[0085] The quantization parameter determination submodule is used to determine the quantization parameters corresponding to each frame of data based on the quantization parameter offset and the standard quantization parameters.
[0086] Furthermore, the device may also include:
[0087] The filtering module is used to construct an offset sequence from the quantization parameter offsets corresponding to each frame of data after determining the quantization parameter offsets corresponding to each frame of data according to the encoding strategy; to perform smoothing filtering on the offset sequence to obtain the filtered offset corresponding to each frame of data; correspondingly, the quantization parameter determination submodule can be used to determine the quantization parameters corresponding to each frame of data based on the filtered offsets and standard quantization parameters.
[0088] As can be seen, the embodiments of this application classify each frame of data based on the luminance feature information corresponding to each frame in the video data to be encoded. Then, based on a preset correspondence, the encoding strategy corresponding to the frame category is determined. Subsequently, the quantization parameters corresponding to each frame of data are determined according to the encoding strategy, and then each frame of data is encoded based on the quantization parameters corresponding to each frame of data to obtain the video encoded data. In this way, by determining the corresponding encoding strategy for each frame of data according to the luminance features of each frame of data, and then determining the quantization parameters used for each frame of data during encoding, the visual perception features of each frame of data are fully considered, and the data is encoded separately. This can improve the efficiency of bit resource utilization while ensuring the quality of video encoding.
[0089] See Figure 3 As shown in the figure, this application discloses an electronic device 20, including a processor 21 and a memory 22; wherein, the memory 22 is used to store a computer program; the processor 21 is used to execute the computer program, the video encoding method disclosed in the foregoing embodiments.
[0090] For details regarding the specific process of the above video encoding method, please refer to the relevant content disclosed in the foregoing embodiments, which will not be repeated here.
[0091] Furthermore, the memory 22, as a carrier for resource storage, can be a read-only memory, random access memory, disk, or optical disk, and the storage method can be temporary storage or permanent storage.
[0092] In addition, the electronic device 20 also includes a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26; wherein, the power supply 23 is used to provide operating voltage for the various hardware devices on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and external devices, and the communication protocol it follows can be any communication protocol applicable to the technical solution of this application, and is not specifically limited here; the input / output interface 25 is used to acquire external input data or output data to the outside world, and its specific interface type can be selected according to specific application needs, and is not specifically limited here.
[0093] Furthermore, embodiments of this application also disclose a computer-readable storage medium for storing a computer program, wherein the computer program, when executed by a processor, implements the video encoding method disclosed in the foregoing embodiments.
[0094] For details regarding the specific process of the above video encoding method, please refer to the relevant content disclosed in the foregoing embodiments, which will not be repeated here.
[0095] Furthermore, embodiments of this application also disclose a computer program product, including a computer program / instructions, which, when executed by a processor, implement the video encoding method disclosed in the foregoing embodiments.
[0096] For details regarding the specific process of the above video encoding method, please refer to the relevant content disclosed in the foregoing embodiments, which will not be repeated here.
[0097] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple; relevant parts can be referred to in the method section.
[0098] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be implemented directly by hardware, a software module executed by a processor, or a combination of both. The software module can be located in random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.
[0099] The above provides a detailed description of a video encoding method, apparatus, device, medium, and product provided in this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.
Claims
1. A video encoding method, characterized in that, include: Obtain the luminance feature information corresponding to each frame of data in the video data to be encoded; The frame category of each frame of data is determined based on the brightness feature information; The encoding strategy corresponding to the frame category is determined based on a preset correspondence. The quantization parameters corresponding to each frame of data are determined according to the encoding strategy described above. Each frame of data is encoded based on the quantization parameters corresponding to each frame of data to obtain video encoded data.
2. The video encoding method according to claim 1, characterized in that, Obtain the luminance feature information corresponding to each frame of data in the video data to be encoded, including: Obtain the maximum and average brightness of each frame in the video data to be encoded.
3. The video encoding method according to claim 2, characterized in that, Determining the frame category of each frame of data based on the brightness feature information includes: The frame category of each frame of data is determined based on the maximum brightness, the average brightness, the preset maximum brightness classification threshold, and the preset average brightness classification threshold.
4. The video encoding method according to claim 3, characterized in that, The preset average brightness classification threshold includes a first threshold and a second threshold, where the first threshold is less than the second threshold. Based on the maximum brightness, the average brightness, and the preset maximum brightness classification threshold and the preset average brightness classification threshold, the frame category of each frame of data is determined, including: For any frame of data, determine whether the average brightness of the frame of data is less than the first threshold. If it is less than the first threshold, determine that the frame of data is a dark frame. If the average brightness of the frame data is greater than or equal to the first threshold, then it is determined whether the average brightness of the frame data is greater than the second threshold and whether the maximum brightness of the frame data is greater than the preset maximum brightness division threshold. If the average brightness of the frame data is greater than the second threshold and the maximum brightness of the frame data is greater than the preset maximum brightness division threshold, then the frame category of the frame data is determined to be a bright frame. If the average brightness of the frame data is less than or equal to the second threshold, and the maximum brightness of the frame data is greater than the preset maximum brightness classification threshold, then the frame category of the frame data is determined to be a mid-gray detail frame. If the average brightness of the frame data is greater than or equal to the second threshold, and the maximum brightness of the frame data is less than or equal to the preset maximum brightness division threshold, then the frame category of the frame data is determined to be a smooth frame.
5. The video encoding method according to claim 4, characterized in that, The quantization parameters for each frame of data are determined according to the encoding strategy, including: The quantization parameter offset for each frame of data is determined according to the encoding strategy described above. The quantization parameters for each frame of data are determined based on the quantization parameter offset and the standard quantization parameters.
6. The video encoding method according to claim 5, characterized in that, After determining the quantization parameter offset corresponding to each frame of data according to the encoding strategy, the method further includes: The quantization parameter offsets corresponding to each frame of data are used to form an offset sequence; The offset sequence is smoothed and filtered to obtain the filtered offset for each frame of data; Accordingly, the quantization parameters for each frame of data are determined based on the quantization parameter offset and the standard quantization parameters, including: The quantization parameters for each frame of data are determined based on the filtered offset and the standard quantization parameters.
7. A video encoding device, characterized in that, include: The luminance feature acquisition module is used to acquire the luminance feature information corresponding to each frame of data in the video data to be encoded; A frame category determination module is used to determine the frame category of each frame of data based on the brightness feature information. The encoding strategy determination module is used to determine the encoding strategy corresponding to the frame category based on a preset correspondence. The quantization parameter determination module is used to determine the quantization parameters corresponding to each frame of data according to the encoding strategy. The video data encoding module is used to encode each frame of data based on the quantization parameters corresponding to each frame of data to obtain video encoded data.
8. An electronic device, characterized in that, Includes memory and processor, wherein: The memory is used to store computer programs; The processor is configured to execute the computer program to implement the video encoding method as described in any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that, Used to store a computer program, wherein the computer program, when executed by a processor, implements the video encoding method as described in any one of claims 1 to 6.
10. A computer program product, characterized in that, Includes a computer program / instruction that, when executed by a processor, implements the video encoding method as described in any one of claims 1 to 6.