A video encoding preprocessing method and device, electronic equipment, storage medium and computer program product
By introducing JND and ROI modules before video encoding, pixel values are adjusted according to human visual characteristics and regions of interest, solving the problem of poor visual effects in existing technologies and achieving better video encoding results.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-04
- Publication Date
- 2026-04-10
AI Technical Summary
Existing video encoders fail to effectively consider human visual characteristics and regions of interest during the encoding process, resulting in poor visual effects after encoding.
By adding a JND preprocessing module and an ROI module, the pixel value correction amount in the target preprocessing mode is determined based on the preprocessing mode parameters, preset encoding parameters, preset region of interest mapping map, and JND value of each pixel of the video frame to be encoded, and the video frame preprocessing is performed.
It improves the visual effect of the encoded video frames, taking into account the characteristics of human vision and the user's area of interest, thereby enhancing video quality.
Smart Images

Figure CN120499379B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the technical field of computers, and particularly relates to a video coding preprocessing method and device, electronic equipment, storage medium and computer program product. BACKGROUND
[0002] In recent years, mobile intelligent terminals have been widely popularized. With the rapid updating of network speed, more and more people prefer lively and intuitive video services. At the same time, people's visual perception is constantly improving, which is reflected in the increasing requirements for video clarity, video fluency and other video quality. Therefore, mainstream video coding standards such as HEVC, AVC, AV1 are more widely used. The main purpose of video coding is to reduce the coding bit rate as much as possible under the condition of ensuring a certain video quality, or to reduce the coding distortion as much as possible under the condition of limiting the mobile coding bit rate, so as to achieve the optimal coding performance.
[0003] In the existing video coding scheme, the video frame is compressed mainly through prediction, transformation, quantization, filtering and other operations. In the rate-distortion optimization process, only the objective quality is optimized, and in the subsequent filtering process, the block effect, ringing effect and other distortions generated in the coding are mainly optimized. Since the human eye visual characteristics and the region of interest (ROI) are not considered in the video coding process, the visual effect after coding may be poor. SUMMARY
[0004] Therefore, the present disclosure provides a video coding preprocessing method and device, electronic equipment, storage medium and computer program product.
[0005] According to an aspect of the present disclosure, a video coding preprocessing method is provided, which includes: determining a target preprocessing mode of a to-be-coded video frame according to a preprocessing mode parameter of the to-be-coded video frame; determining a target pixel value correction amount of each pixel point in the to-be-coded video frame in the target preprocessing mode according to a preset coding parameter, a preset region of interest mapping (ROIMAP) and a preset just noticeable distortion (JND) value of each pixel point in the to-be-coded video frame; and correcting an original pixel value of each pixel point in the to-be-coded video frame according to the target pixel value correction amount of each pixel point in the to-be-coded video frame in the target preprocessing mode, to obtain a preprocessed to-be-coded video frame of the to-be-coded video frame in the target preprocessing mode.
[0006] In a possible implementation, the determining the target pre-processing mode of the to-be-encoded video frame according to the pre-processing mode parameter of the to-be-encoded video frame comprises: determining the target pre-processing mode as a smoothing filter mode when the pre-processing mode parameter is a first pre-processing mode parameter; determining the target pre-processing mode as an enhancement filter mode when the pre-processing mode parameter is a second pre-processing mode parameter; and determining the target pre-processing mode as a smoothing enhancement filter mode when the pre-processing mode parameter is a third pre-processing mode parameter.
[0007] In a possible implementation, the determining the target pixel value correction amount of each pixel point in the to-be-encoded video frame in the target pre-processing mode according to the preset encoding parameter, the preset ROI MAP, and the preset JND value of each pixel point in the to-be-encoded video frame comprises: determining a quantization parameter (QP) value correction amount of each to-be-encoded image block in the to-be-encoded video frame according to the preset ROI MAP; and for any one to-be-encoded image block in the to-be-encoded video frame, determining a target pixel value correction amount of each pixel point in the to-be-encoded image block in the target pre-processing mode according to the preset encoding parameter and the preset JND value of each pixel point in the to-be-encoded image block when the QP value correction amount of the to-be-encoded image block is less than 0.
[0008] In a possible implementation, the target pre-processing mode is a smoothing filter mode, and the determining the target pixel value correction amount of each pixel point in the to-be-encoded image block in the target pre-processing mode according to the preset encoding parameter and the preset JND value of each pixel point in the to-be-encoded image block when the QP value correction amount of the to-be-encoded image block is less than 0 comprises: determining a smoothing-filtered pixel value correction amount of each pixel point in the to-be-encoded image block; determining a pixel value correction amount threshold of each pixel point in the to-be-encoded image block in the smoothing filter mode according to the preset encoding parameter and the preset JND value of each pixel point in the to-be-encoded image block; and for any one pixel point in the to-be-encoded image block, determining a target pixel value correction amount of the pixel point in the smoothing filter mode according to the smoothing-filtered pixel value correction amount of the pixel point and the pixel value correction amount threshold of the pixel point in the smoothing filter mode.
[0009] In a possible implementation, the preset encoding parameter includes a frame type and a frame-level QP value; and the determining, according to the preset encoding parameter and the preset JND value of each pixel point in the to-be-encoded image block, of a pixel value correction threshold of each pixel point in the to-be-encoded image block in the target pre-processing mode includes: in a case where the frame type of the to-be-encoded video frame is a non-I frame and the frame-level QP value of the to-be-encoded video frame is less than or equal to a QP threshold, determining, according to the preset JND value of each pixel point in the to-be-encoded image block and a first JND value correction, of the pixel value correction threshold of the pixel point in the smoothing filtering mode; and in a case where the frame type of the to-be-encoded video frame is a non-I frame and the frame-level QP value of the to-be-encoded video frame is greater than the QP threshold, determining, according to the preset JND value of each pixel point in the to-be-encoded image block and a second JND value correction, of the pixel value correction threshold of the pixel point in the smoothing filtering mode, where the second JND value correction is greater than the first JND value correction.
[0010] In a possible implementation, the target pre-processing mode is an enhancement filtering mode; and the determining, according to the preset encoding parameter and the preset JND value of each pixel point in the to-be-encoded image block, of a target pixel value correction of each pixel point in the to-be-encoded image block in the target pre-processing mode in a case where the QP value correction of the to-be-encoded image block is less than 0 includes: determining an enhancement filtering post-pixel value correction of each pixel point in the to-be-encoded image block; determining a filtering weighting value of each pixel point in the to-be-encoded image block according to the preset encoding parameter; determining a pixel value correction threshold of each pixel point in the to-be-encoded image block in the enhancement filtering mode according to the preset JND value of each pixel point in the to-be-encoded image block and a third JND value correction; and determining, for any one pixel point in the to-be-encoded image block, the target pixel value correction of the pixel point in the enhancement filtering mode according to the enhancement filtering post-pixel value correction of the pixel point, the pixel value correction threshold of the pixel point in the enhancement filtering mode, and the filtering weighting value of the pixel point.
[0011] In a possible implementation, the preset encoding parameters include a frame type and a frame-level QP value; and the determining of the filter weighting value of each pixel point in the to-be-encoded image block according to the preset encoding parameters includes: in a case where the frame type of the to-be-encoded video frame is an I frame and the frame-level QP value of the to-be-encoded video frame is less than or equal to a QP threshold, determining a first filter weighting value corresponding to each pixel point in the to-be-encoded image block; in a case where the frame type of the to-be-encoded video frame is an I frame and the frame-level QP value of the to-be-encoded video frame is greater than the QP threshold, determining a second filter weighting value corresponding to each pixel point in the to-be-encoded image block, wherein the second filter weighting value is greater than the first filter weighting value; in a case where the frame type of the to-be-encoded video frame is a non-I frame and the frame-level QP value of the to-be-encoded video frame is less than or equal to the QP threshold, determining a third filter weighting value corresponding to each pixel point in the to-be-encoded image block; and in a case where the frame type of the to-be-encoded video frame is a non-I frame and the frame-level QP value of the to-be-encoded video frame is greater than the QP threshold, determining a fourth filter weighting value corresponding to each pixel point in the to-be-encoded image block, wherein the fourth filter weighting value is greater than the third filter weighting value.
[0012] In a possible implementation, the target pre-processing mode is a smoothing and enhancement filter mode; and the determining of the target pixel value correction amount of each pixel point in the to-be-encoded video frame in the target pre-processing mode according to the preset encoding parameters of the to-be-encoded video frame, the preset ROI MAP and the preset JND value of each pixel point in the to-be-encoded video frame includes: determining a target pixel value correction amount of each pixel point in the to-be-encoded video frame in a smoothing filter mode according to the preset encoding parameters of the to-be-encoded video frame, the preset ROI MAP and the preset JND value of each pixel point in the to-be-encoded video frame; determining a target pixel value correction amount of each pixel point in the to-be-encoded video frame in an enhancement filter mode according to the preset encoding parameters of the to-be-encoded video frame, the preset ROI MAP and the preset JND value of each pixel point in the to-be-encoded video frame; for any one pixel point in the to-be-encoded video frame, in a case where the target pixel value correction amount of the pixel point in the smoothing filter mode and the target pixel value correction amount of the pixel point in the enhancement filter mode have the same sign, determining the target pixel value correction amount with a smaller value of the two target pixel value correction amounts as the target pixel value correction amount of the pixel point in the smoothing and enhancement filter mode; and for any one pixel point in the to-be-encoded video frame, in a case where the target pixel value correction amount of the pixel point in the smoothing filter mode and the target pixel value correction amount of the pixel point in the enhancement filter mode have different signs, determining the sum of the two target pixel value correction amounts as the target pixel value correction amount of the pixel point in the smoothing and enhancement filter mode.
[0013] According to another aspect of the present disclosure, a video encoding pre-processing apparatus is provided, comprising: a first determining module configured to determine a target pre-processing mode of a to-be-encoded video frame according to a pre-processing mode parameter of the to-be-encoded video frame; a second determining module configured to determine a target pixel value correction amount of each pixel point in the to-be-encoded video frame in the target pre-processing mode according to a preset encoding parameter of the to-be-encoded video frame, a preset region of interest mapping (ROIMAP) and a preset just noticeable distortion (JND) value of each pixel point in the to-be-encoded video frame; and a third determining module configured to correct an original pixel value of each pixel point in the to-be-encoded video frame according to the target pixel value correction amount of each pixel point in the to-be-encoded video frame in the target pre-processing mode, to obtain a pre-processed to-be-encoded video frame of the to-be-encoded video frame in the target pre-processing mode.
[0014] According to another aspect of the present disclosure, an electronic device is provided, comprising a memory, a processor and a computer program stored in the memory, wherein the processor executes the computer program to implement the steps of the above method.
[0015] According to another aspect of the present disclosure, a non-volatile computer readable storage medium is provided, which stores a computer program, wherein the computer program is executed by a processor to implement the steps of the above method.
[0016] According to another aspect of the present disclosure, a computer program product is provided, comprising a computer program, or a non-volatile computer readable storage medium carrying the computer program, wherein the computer program is executed by a processor to implement the steps of the above method.
[0017] In the video encoding preprocessing method according to the embodiments of the present disclosure, according to the preprocessing mode parameter of the to-be-encoded video frame, the target preprocessing mode of the to-be-encoded video frame is determined, and then, the human visual characteristics and the user interested region are taken into account, according to the preset encoding parameter of the to-be-encoded video frame, the preset ROI MAP, and the JND value of each pixel point in the to-be-encoded video frame, the target pixel value correction amount of each pixel point in the to-be-encoded video frame in the target preprocessing mode is determined, so that the original pixel value of each pixel point in the to-be-encoded video frame can be further corrected, and the to-be-encoded video frame after preprocessing in the target preprocessing mode is obtained, thereby effectively improving the visual effect after the to-be-encoded video frame after preprocessing is encoded.
[0018] Other features and aspects of the present disclosure will become apparent from the following detailed description of exemplary embodiments, taken in conjunction with the accompanying drawings. BRIEF DESCRIPTION OF DRAWINGS
[0019] The accompanying drawings, which are incorporated in and constitute a part of the specification, illustrate exemplary embodiments, features, and aspects of the present disclosure and serve to explain the principles of the present disclosure.
[0020] Figure 1 A schematic diagram showing time layers of different video frames in a GOP according to an embodiment of the present disclosure is shown.
[0021] Figure 2 A flowchart of a video encoding preprocessing method according to an embodiment of the present disclosure is shown.
[0022] Figure 3 A schematic diagram showing specific allocation content of 1 byte ROI information transmitted by each 8x8 image block according to an embodiment of the present disclosure is shown.
[0023] Figure 4 A schematic diagram of a mean filter according to an embodiment of the present disclosure is shown.
[0024] Figure 5 A schematic diagram of a low-pass filter according to an embodiment of the present disclosure is shown.
[0025] Figure 6 A block diagram of a video encoding preprocessing apparatus according to an embodiment of the present disclosure is shown.
[0026] Figure 7 A block diagram of an electronic device according to an embodiment of the present disclosure is shown. DETAILED DESCRIPTION
[0027] Various exemplary embodiments, features, and aspects of the present disclosure will be described in detail below with reference to the accompanying drawings. The same reference numbers in different drawings represent the same or similar elements / functionality. Although various aspects of embodiments are illustrated in the drawings, the drawings are not necessarily drawn to scale unless specifically noted.
[0028] As used herein, the terms "comprise", "comprising", "have", "having", "include", "including", "contain", "containing", or variants thereof are open-ended, and include one or more stated features, integers, elements, steps, components or functions but do not preclude the presence or addition of one or more other features, integers, elements, steps, components, functions or groups thereof.
[0029] When an element is referred to as being "connected", "coupled", "responsive", or "in communication" to another element, it can be directly connected, coupled, responsive, or in communication to the other element, or intervening elements can be present.
[0030] Although the terms first, second, third, etc. can be used herein to describe various elements / operations, these elements / operations should not be limited by these terms. These terms are only used to distinguish one element / operation from another element / operation. Thus, a first element / operation in some embodiments can be termed a second element / operation in other embodiments without departing from the teachings of the present inventive concept.
[0031] The word "exemplary" is used herein to mean "serving as an example, instance, or illustration." Any implementation described herein as "exemplary" is not necessarily to be construed as preferred or advantageous over other implementations.
[0032] In addition, for the purpose of convenience and brevity, detailed descriptions of well-known functions and structures incorporated herein can not be described in detail. It should be apparent that the present disclosure can be practiced without such specific details.
[0033] Quantization parameter (QP) in video coding is an important factor to determine the quantization effect. The smaller the QP value is, the smaller the quantization step is, and the smaller the distortion is. The larger the QP value is, the larger the quantization step is, and the more obvious the distortion is.
[0034] Temporal Layer (TL) is a concept in video coding. According to the position of a current video frame in a current Group of Pictures (GOP), the temporal layer in which the current video frame is located can be determined. Figure 1A diagram showing time layers of different video frames in a GOP according to an embodiment of the present disclosure. According to the time interval between a video frame and its reference video frame, the plurality of video frames in the GOP are divided into different time layers. As shown in the figure, video frames of the same color are in the same time layer. The first to eighth video frames belong to a GOP, and the zeroth video frame (zeroth frame in the GOP) does not belong to any GOP. The zeroth video frame and the eighth video frame (eighth frame in the GOP) are key frames, and the time layer is the lowest, which is set to TL=0. Non-key frames are divided into three layers, the fourth video frame (fourth frame in the GOP) is in the TL=1 layer, the second video frame (second frame in the GOP) and the sixth video frame (sixth frame in the GOP) are in the TL=2 layer, and the first, third, fifth, and seventh video frames (first, third, fifth, and seventh frames in the GOP) are in the TL=3 layer. Figure 1 Figure 1 Figure 1 Figure 1 Figure 1 Figure 1 Figure 1
[0035] ROI refers to a certain region or object in an image that is of particular interest, and these regions usually include key information that is intended to be identified, analyzed, or processed.
[0036] In the field of images and videos, Just Noticeable Distortion (JND) models are effectively utilized to reduce visual redundancy. The JND model is a theoretical framework for quantifying the perception threshold based on the characteristics of the human visual system (HVS), and its core goal is to predict the minimum perceptible difference in image or video distortion by mathematical modeling, i.e., assuming that there is a threshold effect in the perception of image distortion by the human eye, i.e., when the degree of distortion is below the threshold, the human eye cannot perceive the difference, thereby providing a scientific basis for image processing, video encoding, and quality assessment, such as objective quality evaluation, image super-resolution, image segmentation, etc.
[0037] In existing video encoder schemes, video frames are compressed mainly through prediction, transformation, quantization, filtering, etc. In the rate-distortion optimization process, only objective quality optimization is considered, and in the subsequent filtering process, optimization is mainly performed for block artifacts, ringing artifacts, and other distortions generated during encoding. Since the human visual characteristics and the region of interest of the human eye are not considered in the video encoding process, the visual effect after encoding may be poor.
[0038] Embodiments of the present disclosure provide a video encoding preprocessing method, which can consider both the human visual characteristics and the region of interest of the user during the video encoding preprocessing process. The video encoding method of the embodiments of the present disclosure is described in detail below.
[0039] Figure 2 A flow chart of a video encoding pre-processing method according to an embodiment of the present disclosure is shown. As shown in Figure 2 The method comprises:
[0040] In step S21, a target pre-processing mode of the video frame to be encoded is determined according to a pre-processing mode parameter of the video frame to be encoded.
[0041] A JND pre-processing module is added in the prior art video encoding process, which is used to pre-process the video frame to be encoded.
[0042] The JND pre-processing module can determine a target pre-processing mode of the video frame to be encoded according to a pre-processing mode parameter prefitter_mode of the video frame to be encoded. One video frame to be encoded uses one target pre-processing mode.
[0043] In one possible implementation, determining the target pre-processing mode of the video frame to be encoded according to the pre-processing mode parameter of the video frame to be encoded comprises: determining the target pre-processing mode as a smoothing filter mode when the pre-processing mode parameter is a first pre-processing mode parameter; determining the target pre-processing mode as an enhancement filter mode when the pre-processing mode parameter is a second pre-processing mode parameter; and determining the target pre-processing mode as a smoothing enhancement filter mode when the pre-processing mode parameter is a third pre-processing mode parameter.
[0044] When the pre-processing mode parameter of the video frame to be encoded is a first pre-processing mode parameter prefitter_mode=0, the target pre-processing mode of the video frame to be encoded is determined as a smoothing filter mode.
[0045] When the pre-processing mode parameter of the video frame to be encoded is a second pre-processing mode parameter prefitter_mode=1, the target pre-processing mode of the video frame to be encoded is determined as an enhancement filter mode.
[0046] When the pre-processing mode parameter of the video frame to be encoded is a third pre-processing mode parameter prefitter_mode=2, the target pre-processing mode of the video frame to be encoded is determined as a smoothing enhancement filter mode.
[0047] In step S22, a target pixel value correction amount of each pixel point in the video frame to be encoded in the target pre-processing mode is determined according to a preset encoding parameter of the video frame to be encoded, a preset ROI MAP, and a preset JND value of each pixel point in the video frame to be encoded.
[0048] A ROI module is added in the prior art video encoding process, which is used to consider a user interested region in the video encoding process, i.e., a preset ROI MAP of each video frame to be encoded.
[0049] Therefore, in the process of pre-processing the to-be-encoded video frame in the target pre-processing mode, for any one pixel point in the to-be-encoded video frame, the preset JND value is used as a threshold, and the preset JND value is processed in combination with the preset encoding parameter and the preset ROIMAP to limit the amplitude of filtering and determine the target pixel value correction amount of the pixel point in the target pre-processing mode.
[0050] In step S23, the original pixel value of each pixel point in the to-be-encoded video frame is corrected according to the target pixel value correction amount of each pixel point in the to-be-encoded video frame in the target pre-processing mode, to obtain the pre-processed to-be-encoded video frame of the to-be-encoded video frame in the target pre-processing mode.
[0051] For any one pixel point in the to-be-encoded video frame, after the target pixel value correction amount of the pixel point in the target pre-processing mode is determined, the original pixel value of the pixel point can be corrected. After the original pixel value of each pixel point in the to-be-encoded video frame is corrected, the pre-processed to-be-encoded video frame is obtained.
[0052] In the video encoding pre-processing method of the embodiments of the present disclosure, according to the pre-processing mode parameter of the to-be-encoded video frame, the target pre-processing mode of the to-be-encoded video frame is determined, and then the human visual characteristics and the user's interested region are considered, and according to the preset encoding parameter, the preset ROIMAP, and the JND value of each pixel point in the to-be-encoded video frame, the target pixel value correction amount of each pixel point in the to-be-encoded video frame in the target pre-processing mode is determined, so that the original pixel value of each pixel point in the to-be-encoded video frame can be further corrected to obtain the pre-processed to-be-encoded video frame of the to-be-processed video frame in the target pre-processing mode, thereby effectively improving the visual effect of the pre-processed to-be-encoded video frame.
[0053] In a possible implementation, the target pixel value correction amount of each pixel point in the to-be-encoded video frame in the target pre-processing mode is determined according to the preset encoding parameter, the preset ROIMAP, and the preset JND value of each pixel point in the to-be-encoded video frame, including: determining the QP value correction amount of each to-be-encoded image block in the to-be-encoded video frame according to the preset ROIMAP; for any one to-be-encoded image block in the to-be-encoded video frame, in the case that the QP value correction amount of the to-be-encoded image block is less than 0, the target pixel value correction amount of each pixel point in the to-be-encoded image block in the target pre-processing mode is determined according to the preset encoding parameter and the preset JND value of each pixel point in the to-be-encoded image block.
[0054] Since the video encoding process on the to-be-encoded video frame is performed in units of image blocks, the preset encoding parameters and the preset ROIMAP of the to-be-encoded video frame are also for each to-be-encoded image block in the to-be-encoded video frame.
[0055] In an example, the preset ROIMAP of the to-be-encoded video frame includes 1 byte of ROI information of each 8x8 image block in the to-be-encoded video frame.
[0056] The 8x8 image block unit can achieve more accurate control of ROI information, and the 8x8 image block is smaller than or equal to the size of the to-be-encoded image block in the to-be-encoded video frame, so that only one-pass is needed to perform the preset ROIMAP corresponding to the single to-be-encoded video frame in the video encoding process, that is, the encoding process of the single to-be-encoded video frame can be completed, which is convenient for hardware implementation of the video encoder.
[0057] The specific size of the ROI information of each 8x8 image block in the ROIMAP can be flexibly set according to the actual application scenario, and the present disclosure does not make specific limitations.
[0058] In an example, for some hardware video encoders or real-time encoding scenarios, in order to improve parallel processing efficiency or memory access alignment, there may be certain requirements for the format of the to-be-encoded video frame, for example, the number of macroblocks (16x16 size) in a horizontal row in the to-be-encoded video frame is a multiple of 32. If the number of macroblocks (16x16 size) in a horizontal row in the original video frame is 21 (corresponding to 336 pixels), then the number of macroblocks needs to be padded to 32 (corresponding to 512 pixels) to obtain the to-be-encoded video frame. The number of macroblocks (16x16 size) in a vertical column in the to-be-encoded video frame can be an integer.
[0059] When there are certain requirements for the format of the to-be-encoded video frame, there are also certain requirements for the preset ROIMAP corresponding to the to-be-encoded video frame. Similarly, assuming that the number of macroblocks (16x16 size) in a horizontal row in the preset ROIMAP corresponding to the to-be-encoded video frame is n, then n is padded according to 32. For a vertical column, it can be directly an integer multiple of the macroblock (16x16 size), and if the number of vertical macroblocks (16x16 size) is not an integer, it is padded to an integer.
[0060] In an example, for the preset ROIMAP of the to-be-encoded video frame, each 8x8 image block transmits 1 byte of ROI information; for the padded block, the transmission value is 0, that is, no ROI information is transmitted.
[0061] The preset ROIMAP of the to-be-encoded video frame includes the preset ROIMAP of each to-be-encoded image block in the to-be-encoded video frame.
[0062] The size of the to-be-encoded image block is N x N, and the specific value of N can be flexibly set according to an actual application scenario, for example, N can be 128, 64, 32, 16, 8, etc., and the present disclosure does not make a specific limitation thereto.
[0063] For any one to-be-encoded image block in the to-be-encoded video frame, the preset QP value of the to-be-encoded image block is adjusted based on the time layer where the to-be-encoded video frame is located and the preset ROI MAP corresponding to the to-be-encoded image block, so as to comprehensively adjust the to-be-encoded image block based on the time layer and the ROI region. The preset QP value of the to-be-encoded image block is pre-allocated to the to-be-encoded image block, and the specific allocation manner can refer to related technologies, and the present disclosure does not make a specific limitation thereto.
[0064] Figure 3 A schematic diagram showing the specific allocation content of 1 byte ROI information of each 8 x 8 image block transmission of the embodiment of the present disclosure. As shown in Figure 3 The 1 byte ROI information of each 8 x 8 image block transmission includes 8 bits, wherein the 0th bit is the position indication parameter Is_first_blk_in_frame of the 8 x 8 image block; the 1st-2nd bits are the preset QP adjustment parameter qp_level of the 8 x 8 image block; the 3rd-6th bits are empty, or other parameters are specified in the 7th bit according to an actual application scenario; and the 7th bit is the ROI adaptive prefiltering parameter roi_adaptive_prefilter_flag of the 8 x 8 image block.
[0065] The ROI information of each 8 x 8 image block transmission can also be set in other forms according to an actual application scenario, and the present disclosure does not make a specific limitation thereto. Figure 3
[0066] In an example, according to the position indication parameter of each 8 x 8 image block included in the to-be-encoded image block, it is judged whether the preset ROI MAP of the to-be-encoded image block exists transmission dislocation.
[0067] For any one 8 x 8 image block, in the case that the position indication parameter Is_first_blk_in_frame of the 8 x 8 image block = 0, it is indicated that the to-be-encoded image block where the 8 x 8 image block is located is not the first to-be-encoded image block in the to-be-encoded video frame; in the case that the position indication parameter Is_first_blk_in_frame of the 8 x 8 image block = 1, it is indicated that the to-be-encoded image block where the 8 x 8 image block is located is the first to-be-encoded image block in the to-be-encoded video frame.
[0068] In the video encoding process, for any one to-be-encoded image block in a to-be-encoded video frame, if there is no transmission dislocation in the preset ROI MAP of the to-be-encoded image block, the QP value correction amount of the to-be-encoded image block is determined according to the preset ROI MAP of the to-be-encoded image block.
[0069] For a to-be-encoded image block in a to-be-encoded video frame, the time layer and the ROI MAP are comprehensively considered, and the QP value correction amount of each 8x8 image block in the to-be-encoded image block is determined.
[0070] For any one 8x8 image block, if the preset QP adjustment parameter qp_level of the 8x8 image block is 0, it is indicated that the 8x8 image block is a non-ROI region; if the preset QP adjustment parameter qp_level of the 8x8 image block is 1, it is indicated that the 8x8 image block is an ROI region and is a low QP adjustment level; if the preset QP adjustment parameter qp_level of the 8x8 image block is 2, it is indicated that the 8x8 image block is an ROI region and is a medium QP adjustment level; and if the preset QP adjustment parameter qp_level of the 8x8 image block is 3, it is indicated that the 8x8 image block is an ROI region and is a high QP adjustment level.
[0071] For a GOP in which a to-be-encoded video frame is located, first, the total time layer number TL_num of the GOP is obtained; then, the time layer threshold T of the GOP is calculated as T = TL_num / 2; and the video frames in the GOP at different time layers are time layer classified by using the time layer threshold T, and are divided into: video frames at a time layer TL greater than the time layer threshold T, and video frames at a time layer TL less than or equal to the time layer threshold T.
[0072] The video frames at the time layer TL greater than the time layer threshold T represent video frames at a higher time layer, and are generally non-reference frames or reference frames with a smaller number of reference times; and the video frames at the time layer TL less than or equal to the time layer threshold T represent video frames at a lower time layer, and have an important reference effect on the encoding of subsequent video frames.
[0073] Taking the above Figure 1 For example, the total time layer number TL_num of the GOP is 4, and thus the time layer threshold T of the GOP is T = TL_num / 2 = 2. At this time, the first, third, fifth and seventh video frames at the time layer TL = 3 are video frames at a time layer greater than the time layer threshold T; the first and eighth video frames at the time layer TL = 0, the fourth video frame at the time layer TL = 1, and the second and sixth video frames at the time layer TL = 2 are video frames at a time layer less than or equal to the time layer threshold T.
[0074] The video frame with the time layer TL greater than the time layer threshold T corresponds to the first preset QP adjustment strategy, and the video frame with the time layer TL less than or equal to the time layer threshold T corresponds to the second preset QP adjustment strategy. Since the QP adjustment range of the second preset QP adjustment strategy is greater than that of the first preset QP adjustment strategy, the video frame with a lower time layer is subjected to a greater enhancement processing, and the video frame with a higher time layer is subjected to a weaker enhancement processing.
[0075] In an example, the first preset QP adjustment strategy and the second preset QP adjustment strategy are shown in Table 1 below.
[0076] Table 1
[0077]
[0078]
[0079] For any 8x8 image block in the to-be-encoded image block, the QP value correction amount delta_qp_8x8 of the 8x8 image block is determined according to the relationship between the time layer TL of the to-be-encoded video frame and the time layer threshold T, the preset QP adjustment parameter qp_level of the 8x8 image block, and the preset QP value qp of the to-be-encoded image block, in combination with Table 1 above.
[0080] As shown in Table 1 above, for the video frame with a lower time layer (TL≤T), the second preset QP adjustment strategy is adopted to perform a stronger enhancement processing (reduce qp) on the ROI block (the first sub-block with qp_level=1, 2, and 3), and the non-ROI block (the first sub-block with qp_level=0) is not processed (qp remains unchanged); for the video frame with a higher time layer (TL>T), the first preset QP adjustment strategy is adopted to perform a weaker enhancement processing (reduce qp) on the ROI block (the first sub-block with qp_level=1, 2, and 3), and the non-ROI block (the first sub-block with qp_level=0) is subjected to a weakening processing (increase qp). In this way, the ROI region can be enhanced while saving a certain amount of encoding code rate.
[0081] As shown in Table 1, for any time layer, when qp≤x3=19, qp is not adjusted, because the original qp is too small, and the increased bit number is too large and the quality of the image is not improved much, which is not cost-effective.
[0082] As shown in Table 1, for the video frame with a higher time layer (TL>T), when qp>x1=39, the non-ROI block (the first sub-block with qp_level=0) will not be subjected to a weakening processing, because the original qp is too large, and the saved bit number is not large, which is not cost-effective.
[0083] The specific values of x1, x2, x3 in Table 1 and the specific values of each delta_qp_8x8 can be flexibly set to other values according to actual conditions, and the present disclosure does not make specific limitations on this.
[0084] After determining the QP value correction amount delta_qp_8x8 of each 8x8 image block in the to-be-encoded image block, the QP value correction amount delta_qp of the entire to-be-encoded image block is determined by synthesizing the QP value correction amount delta_qp_8x8 of each 8x8 image block in the to-be-encoded image block.
[0085] For example, the to-be-encoded image block is 32x32 in size. First, the QP value correction amount delta_qp_8x8 of each 8x8 image block in the 32x32 to-be-encoded image block is determined based on Table 1 by comprehensively considering the temporal layer and the ROI MAP; then, the QP value correction amount delta_qp_8x8 of all 8x8 image blocks in the 32x32 to-be-encoded image block is averaged to obtain the QP value correction amount delta_qp_32x32 of the 32x32 to-be-encoded image block.
[0086] For any one to-be-encoded image block in the to-be-encoded video frame, based on the ROI module, the QP value correction amount of the to-be-encoded image block and the ROI adaptive preprocessing parameter roi_adaptive_prefilter_flag=0 of each 8x8 image block in the to-be-encoded image block are determined.
[0087] For any one 8x8 image block, in the case that the ROI adaptive preprocessing parameter roi_adaptive_prefilter_flag of the 8x8 image block is 0, it is indicated that the 8x8 image block does not need to be preprocessed adaptively; in the case that the ROI adaptive preprocessing parameter roi_adaptive_prefilter_flag of the 8x8 image block is 1, it is indicated that the 8x8 image block needs to be preprocessed adaptively.
[0088] For any one to-be-encoded image block in the to-be-encoded video frame, the target pixel value correction amount of each pixel point in the to-be-encoded image block in the target preprocessing mode is determined.
[0089] In a possible implementation, the target pre-processing mode is a smoothing filter mode; for any one to-be-encoded image block in the to-be-encoded video frame, in a case where a QP value correction amount of the to-be-encoded image block is less than 0, a target pixel value correction amount of each pixel point in the to-be-encoded image block in the target pre-processing mode is determined according to preset encoding parameters and preset JND values of each pixel point in the to-be-encoded image block, including: determining a smoothing-filtered pixel value correction amount of each pixel point in the to-be-encoded image block; determining a pixel value correction amount threshold of each pixel point in the to-be-encoded image block in the smoothing filter mode according to the preset encoding parameters and the preset JND values of each pixel point in the to-be-encoded image block; for any one pixel point in the to-be-encoded image block, a target pixel value correction amount of the pixel point in the smoothing filter mode is determined according to the smoothing-filtered pixel value correction amount of the pixel point and the pixel value correction amount threshold of the pixel point in the smoothing filter mode.
[0090] In a case where the target pre-processing mode is the smoothing filter mode, an input of the JND pre-processing module is an original pixel value org(x, y) of a pixel point, and an output is a pre-processed pixel value jnd_prefilter_out(x, y) of the pixel point in the smoothing filter mode.
[0091] Parameters involved in the pre-processing process include: a smoothing-filtered pixel value mean(x, y) of a pixel point, a smoothing filter coefficient w mean , a preset JND value jnd(x, y) of the pixel point, a pixel value correction amount threshold diff(x, y) of the pixel point in the smoothing filter mode, and a target pixel value correction amount delta(x, y) of the pixel point in the smoothing filter mode.
[0092] In an example, the smoothing filter mode can adopt mean filtering for smoothing processing. Based on a mean filter, mean filtering is performed on the to-be-encoded video frame to obtain a smoothing-filtered pixel value mean(x, y) of a pixel point (x, y) in the to-be-encoded video frame.
[0093] Figure 4 A schematic diagram of a mean filter according to an embodiment of the present disclosure is shown. Based on the 5x5 mean filter shown in Figure 4 The mean filter is used to perform mean filtering on the to-be-encoded video frame to obtain a smoothing-filtered pixel value mean(x, y) of a pixel point (x, y) in the to-be-encoded video frame. k,l org(x+k,y+l)w mean (k,l) / 25.
[0094] In the prior art, mean filtering is directly used in the video encoding preprocessing process. Although the encoding of the video frame to be encoded after mean filtering can save transmission bits, it will cause a lot of detail loss. At the same time, the strength of the preprocessing cannot be adaptively adjusted according to the preset encoding parameters.
[0095] In the embodiments of the present disclosure, the human eye visual characteristics and the user interested region are considered in the video encoding preprocessing process, and the processing idea is as follows:
[0096] (1) Consider the frame type. Since the I frame is generally used as an important reference frame in video encoding, which determines the quality of the subsequent frame, it cannot be smoothed, otherwise it will affect the quality of the subsequent frame, that is, for the I frame, the whole frame is not smoothed. For non-I frames (P frames, B frames), smoothing processing needs to be performed according to the requirements.
[0097] That is, in the case that the frame type of the video frame to be encoded is an I frame, for any pixel point (x, y) in the video frame to be encoded, the pixel value jnd_prefilter_out(x, y) of the pixel point after preprocessing in the smoothing filter mode is the original pixel value org(x, y) of the pixel point.
[0098] (2) Consider the preset QP value. When the frame-level preset QP value of the video frame to be encoded is less than or equal to the QP threshold value, the allowed modification range of the pixel point in the video frame to be encoded during smoothing processing should be large; when the frame-level preset QP value of the video frame to be encoded is greater than the QP threshold value, the allowed modification range of the pixel point in the video frame to be encoded during smoothing processing should be small. This is because when the frame-level preset QP value is large, the texture part itself is blurred, so the amplitude of the smoothing processing should be reduced. The value of the QP threshold value can be flexibly set according to the actual application scene, for example, QP threshold value = 35, which is not limited in the present disclosure.
[0099] (3) Consider the preset JND value. When the preset JND value of the pixel point is large, it means that the human eye cannot easily perceive, so the allowed modification range of the pixel point during smoothing processing should be large; when the preset JND value of the pixel point is small, it means that the human eye pays more attention, so the allowed modification range of the pixel point during smoothing processing should be small.
[0100] In the case that the frame type of the video frame to be encoded is a non-I frame, the pixel value modification threshold of each pixel point in the video frame to be encoded in the smoothing filter mode is determined by comprehensively considering the above (2) and (3) conditions.
[0101] In a possible implementation, the preset encoding parameter includes a frame type and a frame-level QP value; and the determining of the pixel value correction threshold of each pixel point in the to-be-encoded image block in the smoothing filtering mode according to the preset encoding parameter and the preset JND value of each pixel point in the to-be-encoded image block includes: in a case where the frame type of the to-be-encoded video frame is a non-I frame and the frame-level QP value of the to-be-encoded video frame is less than or equal to a QP threshold, the pixel value correction threshold of each pixel point in the to-be-encoded image block in the smoothing filtering mode is determined according to the preset JND value of the pixel point and a first JND value correction; and in a case where the frame type of the to-be-encoded video frame is a non-I frame and the frame-level QP value of the to-be-encoded video frame is greater than the QP threshold, the pixel value correction threshold of each pixel point in the to-be-encoded image block in the smoothing filtering mode is determined according to the preset JND value of the pixel point and a second JND value correction, where the second JND value correction is greater than the first JND value correction.
[0102] Taking a 32x32 to-be-encoded image block as an example, in a case where the frame type of the to-be-encoded video frame is a non-I frame, the ROI adaptive prefilter flag of each 8x8 image block in the 32x32 to-be-encoded image block is 1, and the QP value correction delta_qp_32x32 of the 32x32 to-be-encoded image block is less than 0, that is, in a case where the frame type is a non-I frame, the ROI adaptive prefilter flag is 1, and the QP value correction delta_qp_32x32 is less than 0, the pixel value correction threshold diff(x, y) of any pixel point (x, y) in the 32x32 to-be-encoded image block in the smoothing filtering mode is determined based on the following manner:
[0103] (1) in a case where the preset QP value of the to-be-encoded video frame is less than or equal to a QP threshold (for example, the QP threshold is 35), the filtering amplitude is biased large, the diff(x, y) should be biased large, and therefore, the diff(x, y) = clip(lower threshold, upper threshold, jnd(x, y)-first JND value correction).
[0104] (2) in a case where the preset QP value of the to-be-encoded video frame is greater than the QP threshold (for example, the QP threshold is 35), the filtering amplitude is biased small, the diff(x, y) should be biased small, and therefore, the diff(x, y) = clip(lower threshold, upper threshold, jnd(x, y)-second JND value correction); the second JND value correction is greater than the first JND value correction.
[0105] Wherein, the lower threshold is 0, preventing negative values from leading to incorrect filtering direction; the upper threshold can be flexibly set according to actual application, for example, the upper threshold is 8; the specific values of the first JND value correction amount and the second JND value correction amount can be flexibly set according to actual application, for example, the first JND value correction amount is 4 and the second JND value correction amount is 0. At this time:
[0106] In the case that the preset QP value at the frame level of the to-be-encoded video frame is less than or equal to a QP threshold (for example, the QP threshold is 35), diff(x, y) = clip(0, 8, jnd(x, y) - 4);
[0107] In the case that the preset QP value at the frame level of the to-be-encoded video frame is greater than the QP threshold (for example, the QP threshold is 35), diff(x, y) = clip(0, 8, jnd(x, y) - 8).
[0108] Taking a 32x32 size to-be-encoded image block as an example, in the case that the frame type of the to-be-encoded video frame is a non-I frame, the ROI adaptive pre-processing parameter roi_adaptive_prefilter_flag of each 8x8 image block in the 32x32 to-be-encoded image block is 0, or the QP value correction amount delta_qp_32x32 of the 32x32 to-be-encoded image block is greater than or equal to 0, that is, in the case that the frame type is a non-I frame, roi_adaptive_prefilter_flag is 0 or delta_qp_32x32 is greater than or equal to 0, the pixel value correction amount threshold diff(x, y) of any pixel point (x, y) in the 32x32 to-be-encoded image block in the smoothing filtering mode is determined based on the following manner:
[0109] (1) In the case that the preset QP value at the frame level of the to-be-encoded video frame is less than or equal to a QP threshold (for example, the QP threshold is 35), the filtering amplitude is biased large, and diff(x, y) should be biased large, therefore, diff(x, y) = clip(lower threshold, upper threshold, jnd(x, y) - third JND value correction amount).
[0110] (2) In the case that the preset QP value at the frame level of the to-be-encoded video frame is greater than the QP threshold (for example, the QP threshold is 35), the filtering amplitude is biased small, and diff(x, y) should be biased small, therefore, diff(x, y) = clip(lower threshold, upper threshold, jnd(x, y) - fourth JND value correction amount).
[0111] Wherein, the fourth JND value correction amount is greater than the third JND value correction amount, and the specific values of the third JND value correction amount and the fourth JND value correction amount can be flexibly set according to actual application, for example, the third JND value correction amount is 2 and the fourth JND value correction amount is 4. At this time:
[0112] In case that the preset QP value at frame level of the to-be-encoded video frame is less than or equal to a QP threshold (for example, QP threshold = 35), diff(x, y) = clip(0, 8, jnd(x, y) - 2);
[0113] In case that the preset QP value at frame level of the to-be-encoded video frame is greater than the QP threshold (for example, QP threshold = 35), diff(x, y) = clip(0, 8, jnd(x, y) - 4).
[0114] For any pixel point (x, y) in the to-be-encoded image block, the pixel value change amount mean(x, y) - org(x, y) after smoothing filtering of the pixel point is limited according to the pixel value correction amount threshold diff(x, y) of the pixel point in the smoothing filtering mode, to determine the target pixel value correction amount delta(x, y) = clip(-diff(x, y), diff(x, y), mean(x, y) - org(x, y)) of the pixel point in the smoothing filtering mode, wherein mean(x, y) - org(x, y) is the pixel value correction amount after smoothing filtering of the pixel point (x, y).
[0115] Further, the original pixel value org(x, y) of the pixel point is corrected according to the target pixel value correction amount delta(x, y) of the pixel point in the smoothing filtering mode, to obtain the pre-processed pixel value jnd_prefilter_out(x, y) = org(x, y) + delta(x, y) of the pixel point in the smoothing filtering mode.
[0116] In a possible implementation, the target pre-processing mode is an enhanced filtering mode; for any to-be-encoded image block in the to-be-encoded video frame, in case that the QP value correction amount of the to-be-encoded image block is less than 0, the target pixel value correction amount of each pixel point in the to-be-encoded image block in the target pre-processing mode is determined according to preset encoding parameters and preset JND values of each pixel point in the to-be-encoded image block, including: determining the enhanced pixel value correction amount after filtering of each pixel point in the to-be-encoded image block; determining the filtering weighting value of each pixel point in the to-be-encoded image block according to preset encoding parameters; determining the pixel value correction amount threshold of each pixel point in the to-be-encoded image block in the enhanced filtering mode according to the preset JND value of each pixel point in the to-be-encoded image block and a third JND value correction amount; for any pixel point in the to-be-encoded image block, the target pixel value correction amount of the pixel point in the target pre-processing mode is determined according to the enhanced pixel value correction amount after filtering of the pixel point, the pixel value correction amount threshold of the pixel point in the enhanced filtering mode, and the filtering weighting value of the pixel point.
[0117] In the case that the target pre-processing mode is the enhancement filter mode, the input of the JND pre-processing module is the original pixel value org(x, y) of the pixel point, and the output is the pre-processed pixel value jnd prefilter out(x, y) of the pixel point in the enhancement filter mode.
[0118] The parameters involved in the pre-processing process include: the enhancement filtered pixel value lpf(x, y) of the pixel point, the enhancement filter coefficient w unsharp , the preset JND value jnd(x, y) of the pixel point, the pixel value correction threshold diff(x, y) of the pixel point in the enhancement filter mode, and the target pixel value correction delta(x, y) of the pixel point in the enhancement filter mode.
[0119] In an example, the enhancement filter mode can use unsharp filtering for enhancement processing. Based on the unsharp filter, the unsharp filtering is performed on the to-be-encoded video frame to obtain the enhancement filtered pixel value lpf(x, y) of each pixel point in the to-be-encoded video frame.
[0120] Figure 5 A schematic diagram of a unsharp filter according to an embodiment of the present disclosure is shown. Based on the 5x5 unsharp filter shown in Figure 5 , the unsharp filtering is performed on the to-be-encoded video frame to obtain the enhancement filtered pixel value lpf(x, y) of the pixel point (x, y) in the to-be-encoded video frame. k,l org(x+k, y+l)w unsharp (k, l) / ∑ k,l w lpf (k, l).
[0121] In the prior art, the enhancement filter is directly used in the video encoding pre-processing process. Although the encoding of the enhancement filtered to-be-encoded video frame can improve the visual quality of the video frame, it will cause an increase in the encoding bits. At the same time, the strength of the pre-processing cannot be adaptively adjusted according to the preset encoding parameters.
[0122] In an embodiment of the present disclosure, the human visual characteristics and the user interested region are considered in the video encoding pre-processing process, and the processing idea is as follows:
[0123] (1) Consider the frame type. Since the I frame is generally used as an important reference frame in video encoding, which determines the quality of the subsequent frames, therefore, for the I frame, the allowed modification range of enhancement should be large; for the non-I frame (P frame, B frame), the allowed modification range of enhancement should be small.
[0124] (2) Consider the preset QP value. When the frame-level preset QP value of the to-be-encoded video frame is less than or equal to the QP threshold, the allowed modification range of the pixel enhancement processing in the to-be-encoded video frame should be small; when the frame-level preset QP value of the to-be-encoded video frame is greater than the QP threshold, the allowed modification range of the pixel enhancement processing in the to-be-encoded video frame should be large. This is because when the frame-level preset QP value is small, the transmission overhead required by video coding is large, and the reconstructed video quality is also high. At this time, if the enhancement range is large, more transmission bits will be increased, but the video effect after the video is difficult to improve a lot. Therefore, for the to-be-encoded video frame with a small frame-level preset QP value, the enhancement strength should be appropriately limited. The value of the QP threshold can be flexibly set according to the actual application scene, for example, the QP threshold = 35, which is not limited in the present disclosure.
[0125] (3) Consider the preset JND value. By using the reverse characteristic of JND, when the preset JND value of the pixel point is large, it means that the human eye is not easy to perceive, and the allowed modification range of the pixel enhancement processing should be small; when the preset JND value of the pixel point is small, it means that the human eye pays more attention, and the allowed modification range of the pixel enhancement processing should be large.
[0126] In the case that the frame type of the to-be-encoded video frame is a non-I frame, the pixel value correction threshold of each pixel point in the to-be-encoded video frame in the enhancement filtering mode is determined by comprehensively considering the above (1) (2) (3) conditions.
[0127] In a possible implementation, the preset coding parameters include a frame type and a frame-level QP value; and the filter weighting value of each pixel point in the to-be-encoded image block is determined according to the preset coding parameters, including: in the case that the frame type of the to-be-encoded video frame is an I frame and the frame-level QP value of the to-be-encoded video frame is less than or equal to a QP threshold, determining that each pixel point in the to-be-encoded image block corresponds to a first filter weighting value; in the case that the frame type of the to-be-encoded video frame is an I frame and the frame-level QP value of the to-be-encoded video frame is greater than the QP threshold, determining that each pixel point in the to-be-encoded image block corresponds to a second filter weighting value, wherein the second filter weighting value is greater than the first filter weighting value; in the case that the frame type of the to-be-encoded video frame is a non-I frame and the frame-level QP value of the to-be-encoded video frame is less than or equal to the QP threshold, determining that each pixel point in the to-be-encoded image block corresponds to a third filter weighting value; and in the case that the frame type of the to-be-encoded video frame is a non-I frame and the frame-level QP value of the to-be-encoded video frame is greater than the QP threshold, determining that each pixel point in the to-be-encoded image block corresponds to a fourth filter weighting value, wherein the fourth filter weighting value is greater than the third filter weighting value.
[0128] Taking a 32x32 size to be an example, in a case that the frame type of the to-be-encoded video frame is an I frame, the ROI adaptive pre-processing parameter roi_adaptive_prefilter_flag of each 8x8 image block in the 32x32 to-be-encoded image block is 1, and the QP value correction amount delta_qp_32x32 of the 32x32 to-be-encoded image block is less than 0, that is, in a case that the frame type is an I frame, roi_adaptive_prefilter_flag is 1, and delta_qp_32x32 is less than 0, the filter weighting value λ of any one pixel point (x, y) in the 32x32 to-be-encoded image block in the enhanced filter mode is determined based on the following manner:
[0129] (1) In a case that the preset QP value at the frame level of the to-be-encoded video frame is less than or equal to a QP threshold (for example, the QP threshold is 35), the pixel point corresponds to a first filter weighting value.
[0130] (2) In a case that the preset QP value at the frame level of the to-be-encoded video frame is greater than the QP threshold (for example, the QP threshold is 35), the pixel point corresponds to a second filter weighting value; the second filter weighting value is greater than the first filter weighting value.
[0131] Wherein, the specific values of the first filter weighting value and the second filter weighting value can be flexibly set according to actual application, for example, the first filter weighting value is 0.25, and the second filter weighting value is 0.75. At this time,
[0132] In a case that the preset QP value at the frame level of the to-be-encoded video frame is less than or equal to a QP threshold (for example, the QP threshold is 35), λ = 0.25;
[0133] In a case that the preset QP value at the frame level of the to-be-encoded video frame is greater than the QP threshold (for example, the QP threshold is 35), λ = 0.75.
[0134] Taking a 32x32 size to be an example, in a case that the frame type of the to-be-encoded video frame is an I frame, the ROI adaptive pre-processing parameter roi_adaptive_prefilter_flag of each 8x8 image block in the 32x32 to-be-encoded image block is 1, and the QP value correction amount delta_qp_32x32 of the 32x32 to-be-encoded image block is less than 0, that is, in a case that the frame type is an I frame, roi_adaptive_prefilter_flag is 1, and delta_qp_32x32 is less than 0, the filter weighting value λ of any one pixel point (x, y) in the 32x32 to-be-encoded image block in the enhanced filter mode is determined based on the following manner:
[0135] (1) When the frame-level preset QP value of the video frame to be encoded is less than or equal to the QP threshold (for example, QP threshold = 35), the pixel corresponds to the third filter weighting value.
[0136] (2) When the frame-level preset QP value of the video frame to be encoded is greater than the QP threshold (for example, QP threshold = 35), the pixel corresponds to the fourth filter weighting value; the fourth filter weighting value is greater than the third filter weighting value.
[0137] The specific values of the third and fourth filter weights can be flexibly set according to the actual application. For example, the third filter weight is 0.125, and the fourth filter weight is 0.5. In this case,
[0138] When the frame-level preset QP value of the video frame to be encoded is less than or equal to the QP threshold (e.g., QP threshold = 35), λ = 0.125.
[0139] When the frame-level preset QP value of the video frame to be encoded is greater than the QP threshold (for example, QP threshold = 35), λ = 0.5.
[0140] Taking a 32×32 image block as an example, when the frame type of the video frame to be encoded is I-frame, and the ROI adaptive preprocessing parameter roi_adaptive_prefilter_flag of each 8×8 image block in the 32×32 image block to be encoded is 0, or the QP value correction amount delta_qp_32×32≥0 of the 32×32 image block to be encoded, that is, when the frame type is I-frame, roi_adaptive_prefilter_flag = 0 or delta_qp_32×32≥0, the filter weighting value λ for any pixel (x,y) in the enhanced filtering mode of the 32×32 image block to be encoded is determined based on the following method:
[0141] (1) When the frame-level preset QP value of the video frame to be encoded is less than or equal to the QP threshold (for example, QP threshold = 35), the pixel corresponds to the fifth filter weighting value.
[0142] (2) When the frame-level preset QP value of the video frame to be encoded is greater than the QP threshold (for example, QP threshold = 35), the pixel corresponds to the sixth filter weighting value; the sixth filter weighting value is greater than the fifth filter weighting value.
[0143] The specific values of the fifth and sixth filter weights can be flexibly set according to the actual application. For example, the fifth filter weight can be 0.125, and the sixth filter weight can be 0.5. In this case,
[0144] When the frame-level preset QP value of the video frame to be encoded is less than or equal to the QP threshold (e.g., QP threshold = 35), λ = 0.125.
[0145] In the case that the preset QP value of the frame to be encoded is greater than the QP threshold value (for example, the QP threshold value = 35), λ = 0.5.
[0146] Taking the 32x32 size of the to-be-encoded image block as an example, in the case that the frame type of the to-be-encoded video frame is a non-I frame, the ROI adaptive prefilter flag of each 8x8 image block in the 32x32 to-be-encoded image block is 0, or the QP value correction amount of the 32x32 to-be-encoded image block is greater than 0, that is, the frame type is a non-I frame, the ROI adaptive prefilter flag is 0, or the QP value correction amount of the 32x32 to-be-encoded image block is greater than 0, the filter weighting value λ of any pixel point (x, y) in the 32x32 to-be-encoded image block in the enhanced filter mode is determined based on the following manner:
[0147] (1) In the case that the preset QP value of the frame to be encoded is less than or equal to the QP threshold value (for example, the QP threshold value = 35), the pixel point corresponds to the seventh filter weighting value.
[0148] (2) In the case that the preset QP value of the frame to be encoded is greater than the QP threshold value (for example, the QP threshold value = 35), the pixel point corresponds to the eighth filter weighting value; the eighth filter weighting value is greater than the seventh filter weighting value.
[0149] Wherein, the specific values of the seventh filter weighting value and the eighth filter weighting value can be flexibly set according to actual application, for example, the seventh filter weighting value is 0.0625, and the fourth filter weighting value is 0.25. At this time,
[0150] In the case that the preset QP value of the frame to be encoded is less than or equal to the QP threshold value (for example, the QP threshold value = 35), λ = 0.0625;
[0151] In the case that the preset QP value of the frame to be encoded is greater than the QP threshold value (for example, the QP threshold value = 35), λ = 0.25.
[0152] For any pixel point (x, y) in the to-be-encoded image block, the pixel value correction threshold diff(x, y) in the enhanced filter mode is determined according to the pixel value correction threshold diff(x, y) = clip(threshold lower limit, threshold upper limit, threshold upper limit - jnd(x, y)).
[0153] Wherein, the threshold lower limit is 0, which prevents negative values from causing incorrect filter direction; the threshold upper limit can be flexibly set according to actual application, for example, the threshold upper limit is 8. At this time, the pixel value correction threshold diff(x, y) in the enhanced filter mode of the pixel point (x, y) is diff(x, y) = clip(0, 8, 8 - jnd(x, y)).
[0154] For any one pixel point (x, y) in the to-be-encoded image block, according to a pixel value correction threshold diff(x, y) of the pixel point in the smoothing filter mode, a change amount λ×(org(x, y)-lpf(x, y)) of the enhanced filter post-pixel value of the pixel point is limited, and a target pixel value correction delta(x, y) of the pixel point in the smoothing filter mode is determined as delta(x, y) = clip(-diff(x, y), diff(x, y), λ×(org(x, y)-lpf(x, y))), wherein (org(x, y)-lpf(x, y) is the enhanced filter post-pixel value correction amount of the pixel point (x, y).
[0155] Further, according to the target pixel value correction delta(x, y) of the pixel point in the enhanced filter mode, the original pixel value org(x, y) of the pixel point is corrected to obtain the pre-processing post-pixel value jnd_prefliter_out(x, y) of the pixel point in the enhanced filter mode, jnd_prefliter_out(x, y) = org(x, y) + delta(x, y).
[0156] In a possible implementation, the target pre-processing mode is a smoothing-enhanced filter mode; and the target pixel value correction of each pixel point in the to-be-encoded video frame in the target pre-processing mode is determined according to the preset encoding parameter of the to-be-encoded video frame, the preset ROI MAP, and the preset JND value of each pixel point in the to-be-encoded video frame, including: determining the target pixel value correction of each pixel point in the to-be-encoded video frame in the smoothing filter mode according to the preset encoding parameter of the to-be-encoded video frame, the preset ROI MAP, and the preset JND value of each pixel point in the to-be-encoded video frame; determining the target pixel value correction of each pixel point in the to-be-encoded video frame in the enhanced filter mode according to the preset encoding parameter of the to-be-encoded video frame, the preset ROI MAP, and the preset JND value of each pixel point in the to-be-encoded video frame; for any one pixel point in the to-be-encoded video frame, in a case where the target pixel value correction of the pixel point in the smoothing filter mode and the target pixel value correction of the pixel point in the enhanced filter mode have the same sign, determining the target pixel value correction with a smaller value of the two target pixel value corrections as the target pixel value correction of the pixel point in the smoothing-enhanced filter mode; and in a case where the target pixel value correction of the pixel point in the smoothing filter mode and the target pixel value correction of the pixel point in the enhanced filter mode have different signs, determining the target pixel value correction of the pixel point in the smoothing-enhanced filter mode by summing the two target pixel value corrections.
[0157] In the case that the target pre-processing mode is the smoothing and enhancement filter mode, the input of the JND pre-processing module is the original pixel value org(x, y) of the pixel point, and the output is the pre-processed pixel value jnd_prefilter_out(x, y) of the pixel point in the smoothing and enhancement filter mode.
[0158] The parameters involved in the pre-processing process include: the preset JND value jnd(x, y) of the pixel point, the pixel value lpf(x, y) of the pixel point after the enhancement filter, the pixel value mean(x, y) of the pixel point after the smoothing filter, the pixel value correction threshold diff0(x, y) of the pixel point in the smoothing filter mode, the target pixel value correction delta0(x, y) of the pixel point in the smoothing filter mode, the pixel value correction threshold diff1(x, y) of the pixel point in the enhancement filter mode, the target pixel value correction delta1(x, y) of the pixel point in the enhancement filter mode, and the target pixel value correction delta(x, y) of the pixel point in the smoothing and enhancement filter mode.
[0159] The smoothing and enhancement filter mode combines the advantages of the smoothing filter mode and the enhancement filter mode, achieves the effect of saving the encoding bits of the smoothing filter mode, and also achieves the effect of enhancing the image of the enhancement filter mode. The change amplitude is limited based on the preset JND value, and different processing is performed according to the frame type and the frame-level QP value, so as to improve the visual quality of the human eye while saving the transmission cost.
[0160] The processing idea in the smoothing and enhancement filter mode is as follows:
[0161] If the target pixel value corrections of the smoothing filter mode and the enhancement filter mode are of the same sign, the smaller one is taken as the target pixel value correction in the smoothing and enhancement filter mode; if the target pixel value corrections of the smoothing filter mode and the enhancement filter mode are of different signs, the sum of the two target pixel value corrections is taken as the target pixel value correction in the smoothing and enhancement filter mode.
[0162] Taking a 32x32 size to-be-encoded image block as an example, for any pixel point (x, y) in the 32x32 to-be-encoded block, taking the above-mentioned various value parameters in the smoothing filter mode and the enhancement filter mode as an example (but not limiting other parameter value schemes), the target pixel value correction delta(x, y) of the pixel point (x, y) in the smoothing and enhancement filter mode is determined in the following manner.
[0163] (1) The frame type is I frame, and the frame-level QP value is less than or equal to 35:
[0164] A. Smoothing filter mode:
[0165] The target pixel value correction delta0(x, y) is 0.
[0166] B, Enhancement filter mode:
[0167] pixel value correction threshold diff1(x, y) = clip(0, 8, 8 - jnd(x, y)).
[0168] a) roi_adaptive_prefilter_flag = 1, and delta_qp_32x32 < 0:
[0169] delta1(x, y) = clip(-diff1(x, y), diff1(x, y), 0.25x(org(x, y) - lpf(x, y)));
[0170] b) roi_adaptive_prefilter_flag = 0, or delta_qp_32x32 >= 0:
[0171] delta1(x, y) = clip(-diff1(x, y), diff1(x, y), 0.125x(org(x, y) - lpf(x, y))).
[0172] (2) Frame type is I frame, and frame level QP value > 35:
[0173] A, Smooth filter mode:
[0174] target pixel value correction delta0(x, y) = 0.
[0175] B, Enhancement filter mode:
[0176] pixel value correction threshold diff1(x, y) = clip(0, 8, 8 - jnd(x, y)).
[0177] a) roi_adaptive_prefilter_flag = 1, and delta_qp_32x32 < 0:
[0178] target pixel value correction delta1(x, y) = clip(-diff1(x, y), diff1(x, y), 0.75x(org(x, y) - lpf(x, y)));
[0179] b) roi_adaptive_prefilter_flag = 0, or delta_qp_32x32 >= 0:
[0180] Target pixel value correction amount delta1(x, y) = clip(-diff1(x, y), diff1(x, y), 0.5x(org(x, y)-lpf(x, y))).
[0181] (3) Frame type is non-I frame, and frame level QP value < 35:
[0182] A, roi_adaptive_prefilter_flag = 1, and delta_qp_32x32 < 0
[0183] a) Smooth filtering mode:
[0184] Pixel value correction amount threshold diff0(x, y) = clip(0, 8, jnd(x, y)-4);
[0185] Target pixel value correction amount delta0(x, y) = clip(-diff0(x, y), diff0(x, y), mean(x, y)-org(x, y)).
[0186] b) Enhancement filtering mode:
[0187] Pixel value correction amount threshold diff1(x, y) = clip(0, 8, 8-jnd(x, y)).
[0188] Target pixel value correction amount delta1(x, y) = clip(-diff1(x, y), diff1(x, y), 0.125x(org(x, y)-lpf(x, y))).
[0189] B, roi_adaptive_prefilter_flag = 0, or delta_qp_32x32 >= 0
[0190] a) Smooth filtering mode:
[0191] Pixel value correction amount threshold diff0(x, y) = clip(0, 8, jnd(x, y)-2);
[0192] Target pixel value correction amount delta0(x, y) = clip(-diff0(x, y), diff0(x, y), mean(x, y)-org(x, y)).
[0193] b) Enhancement filtering mode:
[0194] Pixel value correction amount threshold diff1(x, y) = clip(0, 8, 8-jnd(x, y)).
[0195] Target pixel value correction amount delta0(x, y) = clip(-diff0(x, y), diff0(x, y), mean(x, y) - org(x, y)).
[0196] (4) Frame type is non-I frame, and frame level QP value > 35:
[0197] A, roi_adaptive_prefilter_flag = 1, and delta_qp_32x32 < 0
[0198] a) Smooth filtering mode:
[0199] Pixel value correction amount threshold diff0(x, y) = clip(0, 8, jnd(x, y) - 8);
[0200] Target pixel value correction amount delta0(x, y) = clip(-diff0(x, y), diff0(x, y), mean(x, y) - org(x, y)).
[0201] b) Enhancement filtering mode:
[0202] Pixel value correction amount threshold diff1(x, y) = clip(0, 8, 8 - jnd(x, y)).
[0203] Target pixel value correction amount delta1(x, y) = clip(-diff1(x, y), diff1(x, y), 0.5x (org(x, y) - lpf(x, y))).
[0204] B, roi_adaptive_prefilter_flag = 0, or delta_qp_32x32 ≥ 0
[0205] a) Smooth filtering mode:
[0206] Pixel value correction amount threshold diff0(x, y) = clip(0, 8, jnd(x, y) - 4);
[0207] Target pixel value correction amount delta0(x, y) = clip(-diff0(x, y), diff0(x, y), mean(x, y) - org(x, y)).
[0208] b) Enhancement filtering mode:
[0209] Pixel value correction amount threshold diff1(x, y) = clip(0, 8, 8 - jnd(x, y)).
[0210] Target pixel value correction amount delta1(x, y) = clip(-diff1(x, y), diff1(x, y), 0.25 x (org(x, y) - lpf(x, y))).
[0211] The target pixel value correction amount delta(x, y) of the pixel point (x, y) in the smoothing enhancement filter mode is determined according to the following formula:
[0212]
[0213] The preprocessed pixel value jnd_prefilter_out(x, y) of the pixel point in the smoothing enhancement filter mode is org(x, y) + delta(x, y).
[0214] The video encoding preprocessing method of the embodiment of the present disclosure considers the human visual characteristics and the user interested region in the preprocessing process of the to-be-encoded video frame, and the parameters of the preprocessing process are generated according to the preset encoding parameters, different encoding structures are considered, and a video encoding preprocessing method more conducive to encoding is realized.
[0215] In addition, the preprocessing process can include three target preprocessing modes: a smoothing filter mode, an enhancement filter mode, and a smoothing enhancement filter mode. The smoothing filter mode and the enhancement filter mode can be realized by one-time filtering, which has small calculation amount and simple mode and is easy to implement. In the preprocessing process, the frame-level QP value and the frame type are considered, so that the preprocessing process can be optimized according to the characteristics of encoding. In addition, different target preprocessing modes have different processing methods when combined with the ROI. The smoothing processing of the ROI region will weaken, and the increase processing will increase, and the non-ROI region is opposite, which is more targeted.
[0216] Further, by setting the ROI adaptive preprocessing parameter roi_adaptive_prefilter_flag, the adaptive combination switch of the added ROI and JND can adjust whether the preprocessing is adaptive to the ROI information according to the user's needs.
[0217] The video encoding preprocessing method of the embodiment of the present disclosure has no encoding delay and does not introduce too high complexity, and is convenient for hardware video encoder implementation. In addition, the video encoding preprocessing method of the embodiment of the present disclosure does not modify the video encoding standard, and therefore, when used in any video encoder, the video decoder can correctly decode the bitstream to obtain the reconstructed video.
[0218] It can be understood that the above-mentioned various method embodiments mentioned in the present disclosure can be combined with each other to form combined embodiments without deviating from the principle logic. Limited by the length of the present disclosure, the present disclosure will not be described again. Those skilled in the art can understand that in the above-mentioned method of the specific embodiment, the specific execution order of each step should be determined according to its function and possible internal logic.
[0219] In addition, the present disclosure also provides a video coding preprocessing device, an electronic device, a computer readable storage medium, and a program, all of which can be used to implement any one of the video coding preprocessing methods provided by the present disclosure. The corresponding technical solutions and descriptions are described in the method part and are not described again.
[0220] Figure 6 A block diagram of a video coding preprocessing device according to an embodiment of the present disclosure is shown. As shown in Figure 6 The device 60 includes:
[0221] A first determination module 61 is configured to determine a target preprocessing mode of a to-be-coded video frame according to a preprocessing mode parameter of the to-be-coded video frame.
[0222] A second determination module 62 is configured to determine a target pixel value correction amount of each pixel point in the to-be-coded video frame in the target preprocessing mode according to a preset coding parameter of the to-be-coded video frame, a preset ROI MAP, and a preset JND value of each pixel point in the to-be-coded video frame.
[0223] A third determination module 63 is configured to correct an original pixel value of each pixel point in the to-be-coded video frame according to the target pixel value correction amount of each pixel point in the to-be-coded video frame in the target preprocessing mode, to obtain a preprocessed to-be-coded video frame of the to-be-coded video frame in the target preprocessing mode.
[0224] In a possible implementation, the first determination module 61 is specifically configured to:
[0225] In a case where the preprocessing mode parameter is a first preprocessing mode parameter, the target preprocessing mode is determined to be a smoothing filtering mode;
[0226] In a case where the preprocessing mode parameter is a second preprocessing mode parameter, the target preprocessing mode is determined to be an enhancement filtering mode;
[0227] In a case where the preprocessing mode parameter is a third preprocessing mode parameter, the target preprocessing mode is determined to be a smoothing enhancement filtering mode.
[0228] In a possible implementation, the second determination module 62 is specifically configured to:
[0229] According to the preset ROIMAP, a quantization parameter (QP) value correction amount of each to-be-encoded image block in the to-be-encoded video frame is determined;
[0230] For any one to-be-encoded image block in the to-be-encoded video frame, in a case where the QP value correction amount of the to-be-encoded image block is less than 0, a target pixel value correction amount of each pixel point in the to-be-encoded image block in a target pre-processing mode is determined according to preset encoding parameters and preset JND values of each pixel point in the to-be-encoded image block.
[0231] In a possible implementation, the target pre-processing mode is a smoothing filter mode;
[0232] The second determination module 62 is specifically configured to:
[0233] determine a smoothing-filtered pixel value correction amount of each pixel point in the to-be-encoded image block;
[0234] determine a pixel value correction amount threshold of each pixel point in the to-be-encoded image block in the smoothing filter mode according to preset encoding parameters and preset JND values of each pixel point in the to-be-encoded image block;
[0235] For any one pixel point in the to-be-encoded image block, a target pixel value correction amount of the pixel point in the smoothing filter mode is determined according to the smoothing-filtered pixel value correction amount of the pixel point and the pixel value correction amount threshold of the pixel point in the smoothing filter mode.
[0236] In a possible implementation, the preset encoding parameters include a frame type and a frame-level QP value.
[0237] The second determination module 62 is specifically configured to:
[0238] In a case where the frame type of the to-be-encoded video frame is a non-I frame and the frame-level QP value of the to-be-encoded video frame is less than or equal to a QP threshold, a pixel value correction amount threshold of each pixel point in the smoothing filter mode is determined according to preset JND values of the pixel point and the first JND value correction amount.
[0239] In a case where the frame type of the to-be-encoded video frame is a non-I frame and the frame-level QP value of the to-be-encoded video frame is greater than the QP threshold, a pixel value correction amount threshold of each pixel point in the smoothing filter mode is determined according to preset JND values of the pixel point and the second JND value correction amount, where the second JND value correction amount is greater than the first JND value correction amount.
[0240] In a possible implementation, the target pre-processing mode is an enhancement filter mode.
[0241] The second determination module 62 is specifically configured to:
[0242] determine a pixel value correction amount of each pixel in the to-be-encoded image block in the enhancement filter mode according to the preset JND value of each pixel in the to-be-encoded image block and the third JND value correction amount;
[0243] determine a filter weighting value of each pixel in the to-be-encoded image block according to the preset encoding parameter;
[0244] determine a pixel value correction amount threshold of each pixel in the to-be-encoded image block in the enhancement filter mode according to the preset JND value of each pixel in the to-be-encoded image block and the third JND value correction amount;
[0245] for any one pixel in the to-be-encoded image block, determine a target pixel value correction amount of the pixel in the enhancement filter mode according to the pixel value correction amount of the pixel after the enhancement filter, the pixel value correction amount threshold of the pixel in the enhancement filter mode, and the filter weighting value of the pixel.
[0246] In a possible implementation, the preset encoding parameter includes a frame type and a frame-level QP value.
[0247] The second determination module 62 is specifically configured to:
[0248] determine a first filter weighting value corresponding to each pixel in the to-be-encoded image block in a case where the frame type of the to-be-encoded video frame is an I frame and the frame-level QP value of the to-be-encoded video frame is less than or equal to a QP threshold value.
[0249] determine a second filter weighting value corresponding to each pixel in the to-be-encoded image block in a case where the frame type of the to-be-encoded video frame is an I frame and the frame-level QP value of the to-be-encoded video frame is greater than the QP threshold value, wherein the second filter weighting value is greater than the first filter weighting value.
[0250] determine a third filter weighting value corresponding to each pixel in the to-be-encoded image block in a case where the frame type of the to-be-encoded video frame is a non-I frame and the frame-level QP value of the to-be-encoded video frame is less than or equal to the QP threshold value.
[0251] determine a fourth filter weighting value corresponding to each pixel in the to-be-encoded image block in a case where the frame type of the to-be-encoded video frame is a non-I frame and the frame-level QP value of the to-be-encoded video frame is greater than the QP threshold value, wherein the fourth filter weighting value is greater than the third filter weighting value.
[0252] In a possible implementation, the target pre-processing mode is a smoothing enhancement filter mode.
[0253] The second determination module 62 is specifically configured to:
[0254] According to the preset coding parameter of the to-be-encoded video frame, the preset ROI MAP, and the preset JND value of each pixel point in the to-be-encoded video frame, a target pixel value correction amount of each pixel point in the to-be-encoded video frame in the smoothing filtering mode is determined.
[0255] According to the preset coding parameter of the to-be-encoded video frame, the preset ROI MAP, and the preset JND value of each pixel point in the to-be-encoded video frame, a target pixel value correction amount of each pixel point in the to-be-encoded video frame in the enhancement filtering mode is determined.
[0256] For any one pixel point in the to-be-encoded video frame, in a case where the target pixel value correction amount of the pixel point in the smoothing filtering mode and the target pixel value correction amount of the pixel point in the enhancement filtering mode have the same sign, a target pixel value correction amount with a smaller value between the two target pixel value correction amounts is determined as the target pixel value correction amount of the pixel point in the smoothing enhancement filtering mode.
[0257] For any one pixel point in the to-be-encoded video frame, in a case where the target pixel value correction amount of the pixel point in the smoothing filtering mode and the target pixel value correction amount of the pixel point in the enhancement filtering mode have different signs, a sum of the two target pixel value correction amounts is determined as the target pixel value correction amount of the pixel point in the smoothing enhancement filtering mode.
[0258] In some embodiments, the apparatus provided by the embodiments of the present disclosure has functions or includes modules that can be used to execute the methods described in the above method embodiments, and the specific implementation can be referred to the description of the above method embodiments. For briefness, details are not described herein.
[0259] The embodiments of the present disclosure further provide an electronic device, including a memory, a processor, and a computer program stored in the memory, and the processor executes the computer program to implement the steps of the above method.
[0260] The embodiments of the present disclosure further provide a non-volatile computer readable storage medium having a computer program stored thereon, and the computer program is executed by a processor to implement the steps of the above method.
[0261] The embodiments of the present disclosure further provide a computer program product, including a computer program or a non-volatile computer readable storage medium carrying the computer program, and the computer program is executed by a processor to implement the steps of the above method.
[0262] Figure 7 A block diagram of an electronic device according to an embodiment of the present disclosure is shown. Referring to Figure 7 , the apparatus 1900 can be provided as a server or a terminal device. Referring to Figure 7The apparatus 1900 includes a processing assembly 1922, which is further comprised of one or more processors, and memory resources represented by the memory 1932 for storing instructions, such as an application program, executable by the processing assembly 1922. The application programs stored in the memory 1932 can include one or more modules each corresponding to a set of instructions. In addition, the processing assembly 1922 is configured to execute the instructions to perform the methods described above.
[0263] The apparatus 1900 can also include a power supply assembly 1926 configured to perform power management of the apparatus 1900, a wired or wireless network interface 1950 configured to connect the apparatus 1900 to a network, and an input output interface 1958 (I / O interface). The apparatus 1900 can operate based on an operating system stored in the memory 1932, such as Windows Server TM , MacOS X TM , Unix TM , Linux TM , FreeBSD TM or the like.
[0264] In exemplary embodiments, there is also provided a non-transitory computer readable storage medium, such as the memory 1932 comprising computer program instructions executable by the processing assembly 1922 of the apparatus 1900 to perform the methods described above.
[0265] The computer readable storage medium can be a tangible device that can retain and store instructions for execution by an instruction execution device. The computer readable storage medium can be, for example, but is not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer readable storage medium include the following: a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanically encoded device such as punch-cards or punched tape, as well as any suitable combination of the foregoing. A computer readable storage medium, as used herein, is not to be construed as being transitory signals per se, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a waveguide or other transmission media (e.g., light pulses passing through a fiber-optic cable), or electrical signals transmitted through a wire.
[0266] The computer program (or computer readable program instructions) described herein can be downloaded from a computer readable storage medium to various computing / processing devices by way of a network, e.g., the Internet, a local area network, a wide area network, and / or a wireless network. The network can include copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers and / or edge servers. A network adapter card or network interface in each computing / processing device receives computer readable program instructions from the network and forwards the computer readable program instructions for storage in a computer readable storage medium within the respective computing / processing device.
[0267] Computer readable program instructions for carrying out operations of the present disclosure can be assembler instructions, instruction-set-architecture (ISA) instructions, machine instructions, machine dependent instructions, microcode, firmware instructions, state-setting data, or either source code or object code written in any combination of one or more programming languages, including an object oriented programming language such as Smalltalk, C++ or the like, and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The computer readable program instructions can execute entirely on the user's computing / processing device, partly on the user's computing / processing device, as a stand-alone software package, partly on the user's computing / processing device and partly on a remote computing / processing device or entirely on the remote computing / processing device or server. In the latter scenario, the remote computing / processing device can be connected to the user's computing / processing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computing / processing device, for example, through the Internet using an Internet Service Provider. In some embodiments, electronic circuitry including, for example, programmable logic circuitry, field-programmable gate arrays (FPGA), or programmable logic arrays (PLA) can execute the computer readable program instructions by utilizing state information of the computer readable program instructions to personalize the electronic circuitry, in order to perform aspects of the present disclosure.
[0268] The computer readable program instructions can also be loaded onto a computing / processing device, other programmable data processing apparatus, or other device to cause a series of operations to be performed on the computing / processing device, other programmable apparatus or other device to produce a computer implemented process, such that the instructions which execute on the computing / processing device, other programmable apparatus, or other device implement the operations specified in the flow diagrams and / or block diagrams.
[0269] These computer readable program instructions can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks. These computer readable program instructions can also be stored in a computer readable storage medium that can include a non-transitory computer readable storage medium that can be a computer- readable storage medium having no data storage cycles that change state. The instructions can be executed by one or more processors of a computer, to cause a series of operational steps to be performed on the computer to produce a computer-implemented process. The instructions can also cause one or more processors of a computer or other programmable data processing apparatus to
[0270] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer implemented process, such that the instructions which execute on the computer, other programmable data processing apparatus, or other device implement the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0271] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer implemented process, such that the instructions which execute on the computer, other programmable data processing apparatus, or other device implement the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0272] Embodiments of the present disclosure have been described above, and the description is intended to be illustrative of the embodiments and not restrictive of the disclosure. Many modifications and variations of the described embodiments are possible in light of this disclosure without departing from the scope and spirit of the described embodiments. The choice of words in this document is intended to best explain the principles of the embodiments, the practical application, or technical improvement over prior art, or to enable others skilled in the art to understand the embodiments disclosed herein.
Claims
1. A method for video encoding pre-processing, the method comprising: The method comprises the following steps: determining a target pre-processing mode of the to-be-encoded video frame according to a pre-processing mode parameter of the to-be-encoded video frame; determining a target pixel value correction amount of each pixel point in the to-be-encoded video frame in the target pre-processing mode according to preset coding parameters, a preset region of interest mapping (ROIMAP) and preset just noticeable distortion (JND) values of each pixel point in the to-be-encoded video frame, wherein for any one of the pixel points in the to-be-encoded video frame, the preset JND value of the pixel point is processed by combining the preset coding parameters and the preset ROIMAP to limit the filter amplitude, so as to determine the target pixel value correction amount of the pixel point in the target pre-processing mode; correcting the original pixel value of each pixel point in the to-be-encoded video frame according to the target pixel value correction amount of each pixel point in the to-be-encoded video frame in the target pre-processing mode, to obtain a pre-processed to-be-encoded video frame of the to-be-encoded video frame in the target pre-processing mode; determining a target pixel value correction amount of each pixel point in the to-be-encoded video frame in the target pre-processing mode according to preset coding parameters, a preset ROIMAP and preset JND values of each pixel point in the to-be-encoded video frame, comprises: determining a quantization parameter (QP) value correction amount of each to-be-encoded image block in the to-be-encoded video frame according to the preset ROIMAP; for any one of the to-be-encoded image blocks in the to-be-encoded video frame, determining whether to determine a target pixel value correction amount of each pixel point in the to-be-encoded image block in the target pre-processing mode according to the preset coding parameters and the preset JND values of each pixel point in the to-be-encoded image block according to the QP value correction amount of the to-be-encoded image block.
2. The method of claim 1, wherein, The method comprises the following steps: in a case where the pre-processing mode parameter is a first pre-processing mode parameter, determining that the target pre-processing mode is a smoothing filter mode; in a case where the pre-processing mode parameter is a second pre-processing mode parameter, determining that the target pre-processing mode is an enhanced filter mode; in a case where the pre-processing mode parameter is a third pre-processing mode parameter, determining that the target pre-processing mode is a smoothing and enhanced filter mode.
3. The method of claim 1, wherein, The method comprises the following steps: in a case where the QP value correction amount of the to-be-encoded image block is less than 0, determining a target pixel value correction amount of each pixel point in the to-be-encoded image block in the target pre-processing mode according to the preset coding parameters and the preset JND values of each pixel point in the to-be-encoded image block.
4. The method of claim 3, wherein, The target pre-processing mode is a smoothing filter mode. The method comprises the following steps: In the case that the QP value correction amount of the to-be-encoded image block is less than 0, the target pixel value correction amount of each pixel point in the to-be-encoded image block in the target preprocessing mode is determined according to the preset encoding parameter and the preset JND value of each pixel point in the to-be-encoded image block, which comprises the following steps: The smooth-filtered pixel value correction amount of each pixel point in the to-be-encoded image block is determined; The pixel value correction amount threshold of each pixel point in the to-be-encoded image block in the smooth filtering mode is determined according to the preset encoding parameter and the preset JND value of each pixel point in the to-be-encoded image block; 5. The method of claim 4, wherein, The target pixel value correction amount of the pixel point in the smooth filtering mode is determined according to the smooth-filtered pixel value correction amount of the pixel point and the pixel value correction amount threshold of the pixel point in the smooth filtering mode. The preset encoding parameter comprises a frame type and a frame-level QP value; The pixel value correction amount threshold of each pixel point in the to-be-encoded image block in the smooth filtering mode is determined according to the preset encoding parameter and the preset JND value of each pixel point in the to-be-encoded image block, which comprises the following steps: In the case that the frame type of the to-be-encoded video frame is a non-I frame and the frame-level QP value of the to-be-encoded video frame is less than or equal to a QP threshold, the pixel value correction amount threshold of each pixel point in the to-be-encoded image block in the smooth filtering mode is determined according to the preset JND value of each pixel point in the to-be-encoded image block and a first JND value correction amount; 6. The method as claimed in claim 3, wherein, In the case that the frame type of the to-be-encoded video frame is a non-I frame and the frame-level QP value of the to-be-encoded video frame is greater than the QP threshold, the pixel value correction amount threshold of each pixel point in the to-be-encoded image block in the smooth filtering mode is determined according to the preset JND value of each pixel point in the to-be-encoded image block and a second JND value correction amount, wherein the second JND value correction amount is greater than the first JND value correction amount. The target preprocessing mode is an enhanced filtering mode; The target pixel value correction amount of each pixel point in the to-be-encoded image block in the target preprocessing mode is determined according to the preset encoding parameter and the preset JND value of each pixel point in the to-be-encoded image block in the case that the QP value correction amount of the to-be-encoded image block is less than 0, which comprises the following steps: The enhanced-filtered pixel value correction amount of each pixel point in the to-be-encoded image block is determined; The filtering weighting value of each pixel point in the to-be-encoded image block is determined according to the preset encoding parameter; The pixel value correction amount threshold of each pixel point in the to-be-encoded image block in the enhanced filtering mode is determined according to the preset JND value of each pixel point in the to-be-encoded image block and a third JND value correction amount; The target pixel value correction amount of the pixel point in the enhanced filtering mode is determined according to the enhanced-filtered pixel value correction amount of the pixel point, the pixel value correction amount threshold of the pixel point in the enhanced filtering mode and the filtering weighting value of the pixel point.
7. The method of claim 6, wherein, The preset encoding parameter comprises a frame type and a frame-level QP value; The preset encoding parameter comprises a frame type and a frame-level QP value; In a case where the frame type of the to-be-encoded video frame is an I frame and the frame-level QP value of the to-be-encoded video frame is less than or equal to a QP threshold value, a first filter weighting value corresponding to each pixel point in the to-be-encoded image block is determined; In a case where the frame type of the to-be-encoded video frame is an I frame and the frame-level QP value of the to-be-encoded video frame is greater than the QP threshold value, a second filter weighting value corresponding to each pixel point in the to-be-encoded image block is determined, wherein the second filter weighting value is greater than the first filter weighting value; In a case where the frame type of the to-be-encoded video frame is a non-I frame and the frame-level QP value of the to-be-encoded video frame is less than or equal to the QP threshold value, a third filter weighting value corresponding to each pixel point in the to-be-encoded image block is determined; In a case where the frame type of the to-be-encoded video frame is a non-I frame and the frame-level QP value of the to-be-encoded video frame is greater than the QP threshold value, a fourth filter weighting value corresponding to each pixel point in the to-be-encoded image block is determined, wherein the fourth filter weighting value is greater than the third filter weighting value.
8. The method of claim 3, wherein, The target preprocessing mode is a smoothing and enhancement filter mode; According to the preset encoding parameter, the preset ROI MAP and the preset JND value of each pixel point in the to-be-encoded video frame, a target pixel value correction amount of each pixel point in the to-be-encoded video frame in the target preprocessing mode is determined, comprising: According to the preset encoding parameter, the preset ROI MAP and the preset JND value of each pixel point in the to-be-encoded video frame, a target pixel value correction amount of each pixel point in the to-be-encoded video frame in the smoothing filter mode is determined; According to the preset encoding parameter, the preset ROI MAP and the preset JND value of each pixel point in the to-be-encoded video frame, a target pixel value correction amount of each pixel point in the to-be-encoded video frame in the enhancement filter mode is determined; For any one pixel point in the to-be-encoded video frame, in a case where the target pixel value correction amount of the pixel point in the smoothing filter mode and the target pixel value correction amount of the pixel point in the enhancement filter mode have the same sign, the target pixel value correction amount with a smaller value of the two target pixel value correction amounts is determined as the target pixel value correction amount of the pixel point in the smoothing and enhancement filter mode; For any one pixel point in the to-be-encoded video frame, in a case where the target pixel value correction amount of the pixel point in the smoothing filter mode and the target pixel value correction amount of the pixel point in the enhancement filter mode have different signs, the sum of the two target pixel value correction amounts is determined as the target pixel value correction amount of the pixel point in the smoothing and enhancement filter mode.
9. A video encoding pre-processing apparatus, characterized by comprising: The first determining module is configured to determine a target preprocessing mode of the to-be-encoded video frame according to a preprocessing mode parameter of the to-be-encoded video frame; The second determining module is configured to determine a target pixel value correction amount of each pixel point in the to-be-encoded video frame in the target preprocessing mode according to a preset encoding parameter of the to-be-encoded video frame, a preset region of interest mapping (ROIMAP), and a preset just noticeable distortion (JND) value of each pixel point in the to-be-encoded video frame, wherein for any one pixel point in the to-be-encoded video frame, the preset JND value of the pixel point is processed by comprehensively considering the preset encoding parameter and the preset ROIMAP, to limit a filtering amplitude, and the target pixel value correction amount of the pixel point in the target preprocessing mode is determined. The third determining module is configured to correct an original pixel value of each pixel point in the to-be-encoded video frame according to the target pixel value correction amount of each pixel point in the to-be-encoded video frame in the target preprocessing mode, to obtain a preprocessed to-be-encoded video frame of the to-be-encoded video frame in the target preprocessing mode. The second determining module is configured to: determine a QP value correction amount of each to-be-encoded image block in the to-be-encoded video frame according to the preset ROIMAP; and for any one to-be-encoded image block in the to-be-encoded video frame, determine whether to determine a target pixel value correction amount of each pixel point in the to-be-encoded image block in the target preprocessing mode according to the preset encoding parameter and the preset JND value of each pixel point in the to-be-encoded image block according to the QP value correction amount of the to-be-encoded image block.
10. An electronic device comprising a memory, a processor, and a computer program stored on the memory, wherein the computer program, when executed by the processor, is arranged to perform the method of any one of claims 1 to 9. The processor executes the computer program to implement the steps of the method in any one of claims 1 to 8.
11. A non-transitory computer readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to implement the steps of the method in any one of claims 1 to 8.
12. A computer program product comprising a computer program, characterized in that, The computer program is executed by the processor to implement the steps of the method in any one of claims 1 to 8.
Citation Information
Patent Citations
Image processing method and device, electronic equipment and storage medium
CN117764857A