A method, apparatus, device, and medium for HDR video coding based on texture-luminosity-aware masking
By constructing a texture-luminosity joint perception masking model and a nonlinear fusion function, and dynamically adjusting the Lagrange multipliers, the bitrate allocation of the HDR video encoder is optimized, solving the problems of resource waste and insufficient accuracy of the perception model in the existing technology, and achieving more efficient video coding.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- MALANSHAN AUDIO & VIDEO LABORATORY
- Filing Date
- 2026-03-20
- Publication Date
- 2026-05-26
AI Technical Summary
Existing HDR video coding technology fails to accurately combine brightness and texture masking effects, resulting in an overestimation of human eye sensitivity in extremely bright or dark areas, leading to wasted bitrate. Furthermore, the simple linear overlay method cannot reflect the complementary relationship and saturation effect between the two, resulting in insufficient prediction accuracy of the perception model.
A texture-luminance joint perception masking model is constructed. By combining luminance and texture masking components through a nonlinear fusion function and dynamic adjustment of Lagrange multipliers, the model simulates the visual perception law of the human eye and optimizes the bitrate allocation.
While ensuring subjective visual quality, it achieves bitrate savings, improves the quality and efficiency of HDR video encoding, accurately matches the laws of human visual perception, and avoids resource waste.
Smart Images

Figure CN121887988B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of video coding, and in particular to an HDR video coding method, apparatus, device and medium based on texture-luminance perception masking. Background Technology
[0002] Currently, High Dynamic Range (HDR) video, with its wider dynamic range of brightness, richer color gradation, and enhanced detail, has become the mainstream development direction in the video industry. Rate-distortion optimization (RDO), as a crucial step in video coding, aims to minimize distortion under a given bitrate constraint, or minimize the bitrate within an acceptable distortion range. Current mainstream HDR video encoders primarily rely on mathematical distortion metrics such as mean square error and sum square error. These metrics assess coding distortion solely from the perspective of pixel numerical differences, completely detached from the physiological perception characteristics of the human visual system. This leads to severe bitrate waste during coding, meaning a large amount of bitrate is allocated to distortion areas imperceptible to the human eye, while insufficient bitrate is guaranteed for areas sensitive to human vision.
[0003] The perceptual characteristics of the human visual system significantly influence the distortion tolerance of HDR videos, primarily manifested in two masking effects: First, the brightness masking effect: the human eye's contrast sensitivity varies under different average brightness levels, being highest in medium brightness areas and significantly decreasing in extremely bright or dark areas, resulting in higher tolerance for distortion. Second, the texture masking effect: in image areas with complex textures and edges, coded distortion is more easily hidden by texture details, while in flat areas, distortion is easily perceived by the human eye. These two effects are particularly pronounced in HDR videos and are mutually coupled; brightness level and spatial texture together determine the local perceptual sensitivity of the image, and considering only one factor cannot achieve accurate distortion assessment.
[0004] Therefore, existing HDR video coding solutions still have some problems: On the one hand, they only consider the texture masking effect without combining the brightness masking rules in HDR scenes, which leads to an overestimation of human eye sensitivity in extremely bright and dark areas of HDR videos, resulting in excessive bitrate allocation and wasted resources; on the other hand, a few solutions that attempt to combine brightness and texture factors only use a simple linear superposition method to fuse the two types of masking components, which cannot accurately reflect the complementary relationship and saturation effect between the two, resulting in insufficient prediction accuracy of the perception model and difficulty in matching the actual visual perception rules. Summary of the Invention
[0005] In view of this, the purpose of this application is to provide an HDR video coding method, apparatus, device, and medium based on texture-luminosity perceptual masking, which can achieve bitrate savings while ensuring subjective visual quality, thereby improving the quality and efficiency of HDR video coding. The specific solution is as follows:
[0006] In a first aspect, this application provides an HDR video coding method based on texture-luminance-aware masking, applied to an HDR video encoder, comprising:
[0007] The linear brightness average value of the area to be encoded is determined, and the linear brightness average value is converted to the target brightness domain to obtain the first logarithmic value corresponding to the linear brightness average value; the area to be encoded is the area obtained by segmenting the video frame corresponding to the video frame to be encoded, and the elements in the target brightness domain are the logarithmic values corresponding to the brightness.
[0008] The brightness masking component is determined based on the first logarithmic value and the preset function; the preset function is used to simulate the changing pattern of human eye's perception sensitivity to brightness differences in the target brightness domain; the brightness masking component is used to quantify the intensity of the visual masking effect of human eye on video coding distortion under different brightness levels.
[0009] The prediction residual of the image region to be encoded is transformed in the frequency domain to obtain frequency domain coefficients. The frequency domain coefficients are accumulated to obtain an energy value that characterizes the texture complexity of the image region to be encoded. The energy value is then converted into a texture masking component. The texture masking component is used to quantify the intensity of the visual masking effect of the human eye on video encoding distortion under different texture backgrounds.
[0010] The luminance masking component and the texture masking component are mapped to the same preset value range. The luminance masking component and the texture masking component after unifying the value range are fused by a preset nonlinear fusion function to obtain the perceptual factor corresponding to the image area to be encoded. The perceptual factor is used to comprehensively quantify the intensity of the visual masking effect of human eyes on video encoding distortion under the combined effect of luminance and texture.
[0011] The original Lagrange multipliers are adjusted based on the perceptual factors, a rate-distortion cost function is constructed using the adjusted Lagrange multipliers, and video encoding processing is performed on the area to be encoded according to the rate-distortion cost function to obtain the encoded HDR video.
[0012] Optionally, the preset function is a quadratic function that presents a U-shaped curve over the target brightness domain, and its expression is:
[0013] ;
[0014] Where l is the first logarithmic value; The second logarithm corresponds to the preset brightness that the human eye is most sensitive to; Preset positive coefficient; P_light is a preset constant term; P_light is the luminance masking component.
[0015] Optionally, the step of performing a frequency domain transformation on the prediction residual of the image region to be encoded to obtain frequency domain coefficients, performing energy accumulation on the frequency domain coefficients to obtain an energy value characterizing the texture complexity of the image region to be encoded, and converting the energy value into a texture masking component includes:
[0016] A frequency domain transformation method that utilizes difficulty characteristics that meet preset low complexity criteria is used to transform the prediction residual of the image region to be encoded from the spatial domain to the frequency domain to obtain frequency domain coefficients; the frequency domain transformation method includes the Hadamard transform method based on addition and subtraction operations;
[0017] The sum of the absolute values of the frequency domain coefficients is calculated, and the sum of the absolute values of the frequency domain coefficients is used as an energy value characterizing the texture complexity of the image region to be encoded; the larger the energy value, the more complex the texture of the image region to be encoded.
[0018] The energy value is multiplied by a preset weighting coefficient to obtain a texture masking component; the texture masking component is proportional to the energy value.
[0019] Optionally, mapping the luminance masking component and the texture masking component to the same preset numerical range includes:
[0020] Construct a sliding statistics window and obtain the set of luminance masking components and the set of texture masking components corresponding to the image region to be encoded within the current sliding statistics window; wherein, any sliding statistics window corresponds to multiple image regions to be encoded;
[0021] The normalized upper and lower bounds of the brightness masking component set and the normalized upper and lower bounds of the texture masking component set are determined by using preset high quantile points and preset low quantile points.
[0022] Based on the corresponding normalization upper and lower bounds, the luminance masking components in the luminance masking component set and the texture masking components in the texture masking component set are respectively mapped to the same preset value range;
[0023] Specifically, after the sliding statistical window is updated and a new normalized upper and lower bound is redefined, the new normalized upper and lower bounds are subjected to preset time smoothing processing.
[0024] Optionally, the preset nonlinear fusion function is:
[0025] ;
[0026] Where P is the perceptual factor; P'_light is the luminance masking component mapped to a preset numerical range; P'_texture is the texture masking component mapped to a preset numerical range; a is the lower limit of the preset numerical range; and b is the upper limit of the preset numerical range.
[0027] Optionally, adjusting the original Lagrange multipliers based on the perceptual factor and constructing a rate-distortion cost function using the adjusted Lagrange multipliers includes:
[0028] The perception factor is used as a scaling factor and multiplied by the original Lagrange multiplier to obtain the adjusted Lagrange multiplier.
[0029] The adjusted Lagrange multipliers are used to construct a rate-distortion cost function by combining the coding distortion and coding rate.
[0030] Secondly, this application provides an HDR video encoding apparatus based on texture-luminance-aware masking, applied to an HDR video encoder, comprising:
[0031] The first data acquisition module is used to determine the linear brightness average value of the screen area to be encoded, convert the linear brightness average value to the target brightness domain, and obtain the first logarithmic value corresponding to the linear brightness average value; the screen area to be encoded is the screen area obtained by segmenting the video screen corresponding to the video frame to be encoded, and the elements in the target brightness domain are the logarithmic values corresponding to the brightness.
[0032] The second data acquisition module is used to determine the brightness masking component based on the first logarithmic value and the preset function; the preset function is used to simulate the changing pattern of human eye's perception sensitivity to brightness differences in the target brightness domain; and the brightness masking component is used to quantify the intensity of the visual masking effect of human eye on video coding distortion under different brightness levels.
[0033] The third data acquisition module is used to perform frequency domain transformation on the prediction residual of the image region to be encoded to obtain frequency domain coefficients, perform energy accumulation on the frequency domain coefficients to obtain an energy value characterizing the texture complexity of the image region to be encoded, and convert the energy value into a texture masking component; the texture masking component is used to quantify the intensity of the visual masking effect of the human eye on video encoding distortion under different texture backgrounds.
[0034] The data fusion module is used to map the luminance masking component and the texture masking component to the same preset value range, and to fuse the luminance masking component and the texture masking component after unifying the value range through a preset nonlinear fusion function to obtain the perceptual factor corresponding to the image area to be encoded; the perceptual factor is used to comprehensively quantify the intensity of the visual masking effect of human eyes on video encoding distortion under the combined effect of luminance and texture.
[0035] The encoding module is used to adjust the original Lagrange multipliers based on the perceptual factors, construct a rate-distortion cost function using the adjusted Lagrange multipliers, and perform video encoding processing on the area of the image to be encoded according to the rate-distortion cost function to obtain the encoded HDR video.
[0036] Optionally, the third data acquisition module includes:
[0037] The frequency domain transformation unit is used to transform the prediction residual of the image region to be encoded from the spatial domain to the frequency domain using a frequency domain transformation method that meets the preset low complexity judgment conditions, thereby obtaining frequency domain coefficients; the frequency domain transformation method includes the Hadamard transform method based on addition and subtraction operations;
[0038] An energy value determination unit is used to calculate the sum of the absolute values of the frequency domain coefficients and use the sum of the absolute values of the frequency domain coefficients as an energy value characterizing the texture complexity of the image region to be encoded; the larger the energy value, the more complex the texture of the image region to be encoded.
[0039] The data acquisition unit is used to multiply the energy value and a preset weight coefficient to obtain a texture masking component; the texture masking component is proportional to the energy value.
[0040] Thirdly, this application provides an electronic device, comprising:
[0041] Memory, used to store computer programs;
[0042] A processor is configured to execute the computer program to implement the aforementioned HDR video coding method based on texture-luminance perception masking.
[0043] Fourthly, this application provides a computer-readable storage medium for storing a computer program, wherein the computer program, when executed by a processor, implements the aforementioned HDR video coding method based on texture-luminance perception masking.
[0044] In this application, a linear average brightness value of the image region to be encoded is determined, and the linear average brightness value is transformed to a target brightness domain to obtain a first logarithmic value corresponding to the linear average brightness value. The image region to be encoded is the image region obtained by segmenting the video image corresponding to the video frame to be encoded, and the elements in the target brightness domain are the logarithmic values corresponding to brightness. A brightness masking component is determined based on the first logarithmic value and a preset function. The preset function is used to simulate the changing pattern of human eye's perception sensitivity to brightness differences in the target brightness domain, and the brightness masking component is used to quantify the intensity of the visual masking effect of human eye on video encoding distortion under different brightness levels. The prediction residual of the image region to be encoded is subjected to frequency domain transformation to obtain frequency domain coefficients, and the frequency domain coefficients are energy-accumulated to obtain an energy characterizing the texture complexity of the image region to be encoded. The energy value is converted into a texture masking component. The texture masking component is used to quantify the intensity of the visual masking effect of the human eye on video encoding distortion under different texture backgrounds. The luminance masking component and the texture masking component are mapped to the same preset value range. The luminance masking component and the texture masking component after unifying the value range are fused by a preset nonlinear fusion function to obtain the perceptual factor corresponding to the image area to be encoded. The perceptual factor is used to comprehensively quantify the intensity of the visual masking effect of the human eye on video encoding distortion under the combined effect of luminance and texture. Based on the perceptual factor, the original Lagrange multiplier is adjusted, and the rate-distortion cost function is constructed using the adjusted Lagrange multiplier. The image area to be encoded is then processed for video encoding according to the rate-distortion cost function to obtain the encoded HDR video. As can be seen from the above, on the one hand, this application introduces a preset function that simulates the changing law of the human eye's perceptual sensitivity to luminance differences in the target luminance domain to calculate the luminance masking component. This function can accurately reflect the "U-shaped" characteristic of the human eye's contrast sensitivity, which is highest at medium brightness and decreases at both ends of extreme brightness and darkness. By logarithmically processing the linear average brightness of a scene area and substituting it into this function, the visual masking intensity of that area due to the brightness background can be quantified. This step solves a core problem in HDR scenes: in extremely bright and extremely dark areas, the system calculates a higher brightness masking component, meaning the human eye has a higher tolerance, thus proactively reducing the protection priority for these areas in subsequent bitrate allocation, avoiding bitrate waste caused by overestimating sensitivity; on the other hand, a preset nonlinear fusion function is used to combine the brightness masking component and the texture masking component. As long as the masking effect of either brightness or texture is strong, the final perceptual factor will be significantly enhanced; only when both are weak will the perceptual factor be very small, indicating that it needs to be protected.This nonlinear relationship accurately simulates the complex interaction between two masking effects in actual visual perception, where they are both complementary and saturated. Compared to simple linear addition, it can more accurately predict the human eye's true tolerance to composite distortion scenes, thereby improving the prediction accuracy of the perception model. Simultaneously, the perceptual factor is directly integrated into the core decision mechanism of rate-distortion optimization, used to dynamically adjust the Lagrange multiplier. The larger the perceptual factor, the larger the adjusted Lagrange multiplier, and the more the encoder tends to "tolerate distortion and save bit rate" in that region. Conversely, the smaller the perceptual factor, the smaller the adjusted Lagrange multiplier, and the more the encoder tends to "reduce distortion and increase bit rate." Through this mechanism, the encoder's bit rate allocation decision shifts from the traditional minimization of mathematical error to minimizing human-perceived distortion, intelligently redistributing the limited bit rate from areas insensitive to the human eye to areas sensitive to the human eye. This achieves a significant improvement in overall bit rate utilization efficiency while ensuring subjective visual quality. Attached Figure Description
[0045] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of this application. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.
[0046] Figure 1 This is a flowchart of an HDR video coding method based on texture-luminance perception masking disclosed in this application;
[0047] Figure 2 This is a schematic diagram of the structure of an HDR video encoding device based on texture-luminance perception masking disclosed in this application;
[0048] Figure 3 This is a schematic diagram of the structure of an electronic device disclosed in this application. Detailed Implementation
[0049] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0050] Existing HDR video coding techniques still have some problems: on the one hand, considering only the texture masking effect without combining the brightness masking rules in HDR scenes easily leads to resource waste; on the other hand, a few schemes that attempt to combine brightness and texture factors only use a simple linear superposition method to fuse the two types of masking components, which cannot accurately reflect the complementary relationship and saturation effect between the two, and are difficult to match the actual visual perception rules. To this end, this application provides an HDR video coding method based on texture-brightness perceptual masking. By constructing a texture-brightness joint perceptual masking model adapted to HDR video, adopting a nonlinear fusion method, and dynamically adjusting the Lagrange multiplier to optimize the bitrate allocation, bitrate savings are achieved while ensuring subjective visual quality, thus improving the quality and efficiency of HDR video coding.
[0051] See Figure 1 As shown, this application discloses an HDR video coding method based on texture-luminance-aware masking, applied to an HDR video encoder, including:
[0052] Step S11: Determine the linear brightness average value of the area to be encoded, convert the linear brightness average value to the target brightness domain, and obtain the first logarithmic value corresponding to the linear brightness average value; the area to be encoded is the area obtained by segmenting the video frame corresponding to the video frame to be encoded, and the elements in the target brightness domain are the logarithmic values corresponding to the brightness.
[0053] This embodiment constructs a texture-luminosity joint perceptual masking model and uses it to quantify the joint masking effect of the human eye on luminosity and texture in HDR scenes. A perceptual factor is generated to dynamically adjust encoding decisions, thereby guiding the encoder to prioritize bitrate allocation to areas sensitive to human vision, improving compression efficiency while ensuring subjective quality. This texture-luminosity joint perceptual masking model consists of two main components: a luminosity masking component and a texture masking component. The final result is generated by the nonlinear superposition of these two components.
[0054] First, logarithmic processing is performed on the linear brightness average of the area to be encoded:
[0055] ;
[0056] Where L is the average linear brightness of the current image region to be encoded; l is the first logarithm corresponding to the average linear brightness. It is a very small positive number, so avoid taking the logarithm of zero.
[0057] Step S12: Determine the luminance masking component based on the first logarithmic value and the preset function; the preset function is used to simulate the changing pattern of human eye's perception sensitivity to luminance differences in the target luminance domain, and the luminance masking component is used to quantify the intensity of the visual masking effect of human eye on video coding distortion under different luminance conditions.
[0058] In this embodiment, the preset function is a quadratic function that presents a U-shaped curve over the target brightness domain, and its expression is:
[0059] ;
[0060] in, The second logarithm corresponds to the preset brightness that the human eye is most sensitive to; Preset positive coefficient; P_light is a preset constant term; P_light is the luminance masking component.
[0061] Step S13: Perform frequency domain transformation on the prediction residual of the image region to be encoded to obtain frequency domain coefficients, accumulate the energy of the frequency domain coefficients to obtain an energy value characterizing the texture complexity of the image region to be encoded, and convert the energy value into a texture masking component; the texture masking component is used to quantify the intensity of the visual masking effect of the human eye on video encoding distortion under different texture backgrounds.
[0062] In this embodiment, firstly, a frequency domain transformation method that meets the preset low complexity judgment condition can be used to transform the prediction residual of the image region to be encoded from the spatial domain to the frequency domain, obtaining frequency domain coefficients. The frequency domain transformation method includes, but is not limited to, the Hadamard transform method based on addition and subtraction operations. Then, the sum of the absolute values of the frequency domain coefficients can be calculated, and this sum is used as an energy value representing the texture complexity of the image region to be encoded; the larger the energy value, the more complex the texture of the image region to be encoded. Finally, the energy value can be multiplied by a preset weight coefficient to obtain a texture masking component; the texture masking component is proportional to the energy value.
[0063] For example, the formula for calculating texture masking components is:
[0064] ;
[0065] in, The weighting coefficients for adjusting the texture masking effect are: E_texture is the energy value representing the texture complexity of the image region to be encoded; P_texture is the texture masking component. The smaller the P_texture value, the stronger the visual masking effect of video coding distortion on the human eye under the current corresponding brightness, that is, the higher the tolerance of the human eye to video coding distortion under the current corresponding brightness.
[0066] It should be noted that other frequency domain transformation methods can also be used to obtain high-frequency responses in this embodiment, but the Hadamard transform has the advantage of only including addition and subtraction operations, which can significantly reduce complexity.
[0067] Step S14: Map the luminance masking component and the texture masking component to the same preset value range, and fuse the luminance masking component and the texture masking component after unifying the value range through a preset nonlinear fusion function to obtain the perceptual factor corresponding to the image area to be encoded; the perceptual factor is used to comprehensively quantify the intensity of the visual masking effect of human eyes on video encoding distortion under the combined effect of luminance and texture.
[0068] Since the numerical ranges of P_light and P_texture cannot be predetermined and fluctuate with scene changes, direct use would lead to oscillations or instability. Therefore, this embodiment proposes an adaptive normalization mechanism to map arbitrary ranges of P_light and P_texture to a preset stable numerical range, providing robustness against sudden scene changes. Furthermore, common software encoders (x264 / x265 / VVC) have lookahead mechanisms for several frames. For such scenarios, this embodiment employs sliding window statistics.
[0069] Specifically, a sliding statistical window can be constructed first to obtain the sets of luminance masking components and texture masking components corresponding to the image regions to be encoded within the current sliding statistical window; each sliding statistical window corresponds to multiple image regions to be encoded. Then, preset high and low quantiles can be used to determine the normalized upper and lower bounds corresponding to the luminance masking component set, and the normalized upper and lower bounds corresponding to the texture masking component set. Finally, based on the corresponding normalized upper and lower bounds, the luminance masking components in the luminance masking component set and the texture masking components in the texture masking component set are mapped to the same preset numerical range; after the sliding statistical window is updated and new normalized upper and lower bounds are redefined, preset time smoothing processing is performed on the new normalized upper and lower bounds to avoid jitter.
[0070] For example, mapping P_light and P_texture of any range to the stable interval [0,1]:
[0071] First, collect the P_light and P_texture values corresponding to all areas of the image to be encoded within the sliding statistical window. Then, calculate the 99th and 1st quantiles of all P_light values, and use them as the normalization upper bound P_lightmax and normalization lower bound P_lightmin, respectively. The same method is used to obtain P_texturemax and P_texturemin for P_texture. Finally, perform normalization to map P_light and P_texture to the stable interval [0,1].
[0072] P_lightnorm=clip((P_light-P_lightmin) / (P_lightmax-P_lightmin), 0, 1);
[0073] P_texturenorm=clip((P_texture-P_texturemin) / (P_texturemax-P_texturemin), 0, 1).
[0074] Furthermore, the luminance masking component and texture masking component, after unifying the numerical range, can be fused using a preset nonlinear fusion function to obtain the perceptual factor corresponding to the image region to be encoded; wherein, the preset nonlinear fusion function is:
[0075] ;
[0076] P is the perceptual factor; P'_light is the luminance masking component mapped to a preset value range; P'_texture is the texture masking component mapped to a preset value range; a is the lower limit of the preset value range; b is the upper limit of the preset value range.
[0077] As shown above, when P'_light or P'_texture is strong, P will be closer to b, indicating a strong visual masking effect. The human eye has a high tolerance for encoding distortion in this area, allowing for a lower bitrate allocation. If both P'_light and P'_texture are weak, P will be closer to a, indicating a weak visual masking effect. Distortion is easily detected, requiring a higher bitrate allocation. Compared to linear fusion, this method more accurately reflects the coupling relationship between brightness and texture masking, allowing the perceptual factor P to better guide rate-distortion optimization and bitrate allocation in HDR video coding.
[0078] Step S15: Adjust the original Lagrange multipliers based on the perceptual factor, construct a rate-distortion cost function using the adjusted Lagrange multipliers, and perform video encoding processing on the area to be encoded according to the rate-distortion cost function to obtain the encoded HDR video.
[0079] In this embodiment, adjusting the original Lagrange multipliers based on the perceptual factor and constructing the rate-distortion cost function using the adjusted Lagrange multipliers can include: multiplying the perceptual factor as a scaling factor with the original Lagrange multipliers to obtain the adjusted Lagrange multipliers. Then, the adjusted Lagrange multipliers are used, combined with the coding distortion and coding rate, to construct the rate-distortion cost function. For example:
[0080] ;
[0081] ;
[0082] in, These are the original Lagrange multipliers; is the adjusted Lagrange multiplier; J is the rate-distortion cost; D is the coding distortion degree; R is the coding rate.
[0083] It's important to note that when P is greater than 1, it indicates a stronger masking effect in the area to be encoded. This could mean the area is in the extremely bright / dark range of an HDR video, or it might be filled with complex textures and edges, or even both. In this case, the human eye can hardly perceive the slight distortion introduced by encoding, so the encoder can relax distortion control and allocate less bitrate to this area, reducing bitrate consumption without affecting subjective visual quality and avoiding resource waste. When P is less than 1, it indicates a weaker masking effect in the area to be encoded. This typically means the area is in the medium brightness range of an HDR video, and the image is flat with no complex textures. In this case, even the slightest encoding distortion can be clearly perceived by the human eye. Therefore, the encoder must strictly control distortion and actively increase bitrate allocation to this area to reduce distortion and ensure subjective image quality, avoiding significant impact on the viewing experience due to distortion.
[0084] As shown above, this embodiment constructs a proprietary texture-luminance joint perceptual masking model for HDR videos, achieving rate-distortion optimization that highly matches the characteristics of human vision. This solves the problem of inaccurate bitrate allocation in traditional methods in HDR scenes, demonstrating significant advantages. On the other hand, by integrating the luminance masking component and the texture masking component through a nonlinear fusion function, the complementary and saturation effects of the two in actual perception are more accurately reflected, improving the model's prediction accuracy. Simultaneously, the calculated perceptual factors are dynamically applied to the Lagrange multipliers, driving the encoder to intelligently allocate more bitrate to the human eye's sensitive areas, thereby achieving significant bitrate savings while ensuring excellent subjective visual quality. Furthermore, some calculations in the model, such as the Hadamard transform, are compatible with existing coding frameworks, ensuring low complexity and high integrability of the solution.
[0085] See Figure 2As shown in the embodiments, this application also discloses an HDR video encoding device based on texture-luminance perceptual masking, applied to an HDR video encoder, comprising:
[0086] The first data acquisition module 11 is used to determine the average linear brightness of the screen area to be encoded, convert the average linear brightness to a target brightness domain, and obtain the first logarithmic value corresponding to the average linear brightness; the screen area to be encoded is the screen area obtained by segmenting the video screen corresponding to the video frame to be encoded, and the elements in the target brightness domain are the logarithmic values corresponding to the brightness.
[0087] The second data acquisition module 12 is used to determine the brightness masking component based on the first logarithmic value and the preset function; the preset function is used to simulate the changing pattern of human eye's perception sensitivity to brightness differences in the target brightness domain; and the brightness masking component is used to quantify the intensity of the visual masking effect of human eye on video coding distortion under different brightness levels.
[0088] The third data acquisition module 13 is used to perform frequency domain transformation on the prediction residual of the image region to be encoded to obtain frequency domain coefficients, perform energy accumulation on the frequency domain coefficients to obtain an energy value characterizing the texture complexity of the image region to be encoded, and convert the energy value into a texture masking component; the texture masking component is used to quantify the intensity of the visual masking effect of the human eye on video encoding distortion under different texture backgrounds.
[0089] The data fusion module 14 is used to map the luminance masking component and the texture masking component to the same preset value range, and to fuse the luminance masking component and the texture masking component after unifying the value range through a preset nonlinear fusion function to obtain the perceptual factor corresponding to the image area to be encoded; the perceptual factor is used to comprehensively quantify the intensity of the visual masking effect of human eyes on video encoding distortion under the combined effect of luminance and texture.
[0090] The encoding module 15 is used to adjust the original Lagrange multipliers based on the perceptual factor, construct a rate-distortion cost function using the adjusted Lagrange multipliers, and perform video encoding processing on the area of the image to be encoded according to the rate-distortion cost function to obtain the encoded HDR video.
[0091] In some specific embodiments, the preset function is a quadratic function that presents a U-shaped curve over the target brightness domain, and its expression is:
[0092] ;
[0093] Where l is the first logarithmic value; The second logarithm corresponds to the preset brightness that the human eye is most sensitive to; Preset positive coefficient; P_light is a preset constant term; P_light is the luminance masking component.
[0094] In some specific embodiments, the third data acquisition module 13 includes:
[0095] The frequency domain transformation unit is used to transform the prediction residual of the image region to be encoded from the spatial domain to the frequency domain using a frequency domain transformation method that meets the preset low complexity judgment conditions, thereby obtaining frequency domain coefficients; the frequency domain transformation method includes the Hadamard transform method based on addition and subtraction operations;
[0096] An energy value determination unit is used to calculate the sum of the absolute values of the frequency domain coefficients and use the sum of the absolute values of the frequency domain coefficients as an energy value characterizing the texture complexity of the image region to be encoded; the larger the energy value, the more complex the texture of the image region to be encoded.
[0097] The data acquisition unit is used to multiply the energy value and a preset weight coefficient to obtain a texture masking component; the texture masking component is proportional to the energy value.
[0098] In some specific embodiments, the data fusion module 14 includes:
[0099] The data acquisition unit is used to construct a sliding statistics window and acquire the set of luminance masking components and the set of texture masking components corresponding to the image region to be encoded within the current sliding statistics window; wherein, any sliding statistics window corresponds to multiple image regions to be encoded;
[0100] The upper and lower bound determination unit is used to determine the normalized upper bound and normalized lower bound of the brightness masking component set, and the normalized upper bound and normalized lower bound of the texture masking component set, using preset high quantile points and preset low quantile points.
[0101] An interval unification unit is used to map the luminance masking components in the luminance masking component set and the texture masking components in the texture masking component set to the same preset numerical interval according to the corresponding normalization upper and lower bounds.
[0102] Specifically, after the sliding statistical window is updated and a new normalized upper and lower bound is redefined, the new normalized upper and lower bounds are subjected to preset time smoothing processing.
[0103] In some specific implementations, the preset nonlinear fusion function is:
[0104] ;
[0105] Where P is the perceptual factor; P'_light is the luminance masking component mapped to a preset numerical range; P'_texture is the texture masking component mapped to a preset numerical range; a is the lower limit of the preset numerical range; and b is the upper limit of the preset numerical range.
[0106] In some specific embodiments, the encoding module 15 includes:
[0107] The adjustment unit is used to multiply the perception factor as a scaling factor with the original Lagrange multiplier to obtain the adjusted Lagrange multiplier.
[0108] The function construction unit is used to construct a rate-distortion cost function by utilizing the adjusted Lagrange multipliers and combining the coding distortion and coding rate.
[0109] Furthermore, embodiments of this application also disclose an electronic device, Figure 3 This is a structural diagram of an electronic device 20 according to an exemplary embodiment. The content of the diagram should not be construed as limiting the scope of this application.
[0110] Figure 3 This is a schematic diagram of the structure of an electronic device 20 provided in an embodiment of this application. Specifically, the electronic device 20 may include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. The memory 22 stores a computer program, which is loaded and executed by the processor 21 to implement the relevant steps in the texture-luminance-aware masking-based HDR video coding method disclosed in any of the foregoing embodiments. Alternatively, the electronic device 20 in this embodiment may specifically be a computer.
[0111] In this embodiment, the power supply 23 is used to provide operating voltage for each hardware device on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and external devices, and the communication protocol it follows can be any communication protocol applicable to the technical solution of this application, and is not specifically limited here; the input / output interface 25 is used to acquire external input data or output data to the outside world, and its specific interface type can be selected according to specific application needs, and is not specifically limited here.
[0112] In addition, the memory 22, as a carrier for resource storage, can be a read-only memory, random access memory, disk or optical disk, etc. The resources stored thereon can include operating system 221, computer program 222, etc., and the storage method can be temporary storage or permanent storage.
[0113] The operating system 221 is used to manage and control the various hardware devices on the electronic device 20 and the computer program 222, which may be Windows Server, Netware, Unix, Linux, etc. In addition to including a computer program capable of performing the texture-luminance-aware masking-based HDR video coding method disclosed in any of the foregoing embodiments, the computer program 222 may further include computer programs capable of performing other specific tasks.
[0114] Furthermore, this application also discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, it implements the aforementioned HDR video coding method based on texture-luminance perceptual masking. Specific steps of this method can be found in the corresponding content disclosed in the foregoing embodiments, and will not be repeated here.
[0115] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple; relevant parts can be referred to in the method section.
[0116] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0117] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be implemented directly by hardware, a software module executed by a processor, or a combination of both. The software module can be located in random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.
[0118] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0119] The technical solutions provided in this application have been described in detail above. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the methods and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.
Claims
1. An HDR video coding method based on texture-luminance-aware masking, characterized in that, Applied to HDR video encoders, including: The linear brightness average value of the area to be encoded is determined, and the linear brightness average value is converted to the target brightness domain to obtain the first logarithmic value corresponding to the linear brightness average value; the area to be encoded is the area obtained by segmenting the video frame corresponding to the video frame to be encoded, and the elements in the target brightness domain are the logarithmic values corresponding to the brightness. The brightness masking component is determined based on the first logarithmic value and the preset function; the preset function is used to simulate the changing pattern of human eye's perception sensitivity to brightness differences in the target brightness domain; the brightness masking component is used to quantify the intensity of the visual masking effect of human eye on video coding distortion under different brightness levels. The prediction residual of the image region to be encoded is transformed in the frequency domain to obtain frequency domain coefficients. The frequency domain coefficients are accumulated to obtain an energy value that characterizes the texture complexity of the image region to be encoded. The energy value is then converted into a texture masking component. The texture masking component is used to quantify the intensity of the visual masking effect of the human eye on video encoding distortion under different texture backgrounds. The luminance masking component and the texture masking component are mapped to the same preset value range. The luminance masking component and the texture masking component after unifying the value range are fused by a preset nonlinear fusion function to obtain the perceptual factor corresponding to the image area to be encoded. The perceptual factor is used to comprehensively quantify the intensity of the visual masking effect of human eyes on video encoding distortion under the combined effect of luminance and texture. The original Lagrange multipliers are adjusted based on the perceptual factors, a rate-distortion cost function is constructed using the adjusted Lagrange multipliers, and video encoding processing is performed on the area to be encoded according to the rate-distortion cost function to obtain the encoded HDR video. The preset function is a quadratic function that presents a U-shaped curve in the target brightness domain, and its expression is: ; Where l is the first logarithmic value; The second logarithm corresponds to the preset brightness that the human eye is most sensitive to; Preset positive coefficient; P_light is a preset constant term; P_light is the luminance masking component.
2. The HDR video coding method based on texture-luminance perceptual masking according to claim 1, characterized in that, The process of performing a frequency domain transformation on the prediction residual of the image region to be encoded to obtain frequency domain coefficients, accumulating the energy of the frequency domain coefficients to obtain an energy value characterizing the texture complexity of the image region to be encoded, and converting the energy value into a texture masking component includes: A frequency domain transformation method that utilizes difficulty characteristics that meet preset low complexity criteria is used to transform the prediction residual of the image region to be encoded from the spatial domain to the frequency domain to obtain frequency domain coefficients; the frequency domain transformation method includes the Hadamard transform method based on addition and subtraction operations; The sum of the absolute values of the frequency domain coefficients is calculated, and the sum of the absolute values of the frequency domain coefficients is used as an energy value characterizing the texture complexity of the image region to be encoded; the larger the energy value, the more complex the texture of the image region to be encoded. The energy value is multiplied by a preset weighting coefficient to obtain a texture masking component; the texture masking component is proportional to the energy value.
3. The HDR video coding method based on texture-luminance perceptual masking according to claim 1, characterized in that, The step of mapping the luminance masking component and the texture masking component to the same preset numerical range includes: Construct a sliding statistics window and obtain the set of luminance masking components and the set of texture masking components corresponding to the image region to be encoded within the current sliding statistics window; wherein, any sliding statistics window corresponds to multiple image regions to be encoded; The normalized upper and lower bounds of the brightness masking component set and the normalized upper and lower bounds of the texture masking component set are determined by using preset high quantile points and preset low quantile points. Based on the corresponding normalization upper and lower bounds, the luminance masking components in the luminance masking component set and the texture masking components in the texture masking component set are respectively mapped to the same preset value range; Specifically, after the sliding statistical window is updated and a new normalized upper and lower bound is redefined, the new normalized upper and lower bounds are subjected to preset time smoothing processing.
4. The HDR video coding method based on texture-luminance perceptual masking according to claim 1, characterized in that, The preset nonlinear fusion function is: ; Where P is the perceptual factor; P'_light is the luminance masking component mapped to a preset numerical range; P'_texture is the texture masking component mapped to a preset numerical range; a is the lower limit of the preset numerical range; and b is the upper limit of the preset numerical range.
5. The HDR video coding method based on texture-luminance perceptual masking according to claim 1, characterized in that, The step of adjusting the original Lagrange multipliers based on the perceptual factor and constructing a rate-distortion cost function using the adjusted Lagrange multipliers includes: The perception factor is used as a scaling factor and multiplied by the original Lagrange multiplier to obtain the adjusted Lagrange multiplier. The adjusted Lagrange multipliers are used to construct a rate-distortion cost function by combining the coding distortion and coding rate.
6. An HDR video encoding device based on texture-luminance perceptual masking, characterized in that, Applied to HDR video encoders, including: The first data acquisition module is used to determine the linear brightness average value of the screen area to be encoded, convert the linear brightness average value to the target brightness domain, and obtain the first logarithmic value corresponding to the linear brightness average value; the screen area to be encoded is the screen area obtained by segmenting the video screen corresponding to the video frame to be encoded, and the elements in the target brightness domain are the logarithmic values corresponding to the brightness. The second data acquisition module is used to determine the brightness masking component based on the first logarithmic value and the preset function; the preset function is used to simulate the changing pattern of human eye's perception sensitivity to brightness differences in the target brightness domain; and the brightness masking component is used to quantify the intensity of the visual masking effect of human eye on video coding distortion under different brightness levels. The third data acquisition module is used to perform frequency domain transformation on the prediction residual of the image region to be encoded to obtain frequency domain coefficients, perform energy accumulation on the frequency domain coefficients to obtain an energy value characterizing the texture complexity of the image region to be encoded, and convert the energy value into a texture masking component; the texture masking component is used to quantify the intensity of the visual masking effect of the human eye on video encoding distortion under different texture backgrounds. The data fusion module is used to map the luminance masking component and the texture masking component to the same preset value range, and to fuse the luminance masking component and the texture masking component after unifying the value range through a preset nonlinear fusion function to obtain the perceptual factor corresponding to the image area to be encoded; the perceptual factor is used to comprehensively quantify the intensity of the visual masking effect of human eyes on video encoding distortion under the combined effect of luminance and texture. The encoding module is used to adjust the original Lagrange multipliers based on the perceptual factors, construct a rate-distortion cost function using the adjusted Lagrange multipliers, and perform video encoding processing on the area of the image to be encoded according to the rate-distortion cost function to obtain the encoded HDR video. The preset function is a quadratic function that presents a U-shaped curve in the target brightness domain, and its expression is: ; Where l is the first logarithmic value; The second logarithm corresponds to the preset brightness that the human eye is most sensitive to; Preset positive coefficient; P_light is a preset constant term; P_light is the luminance masking component.
7. The HDR video encoding apparatus based on texture-luminance perceptual masking according to claim 6, characterized in that, The third data acquisition module includes: The frequency domain transformation unit is used to transform the prediction residual of the image region to be encoded from the spatial domain to the frequency domain using a frequency domain transformation method that meets the preset low complexity judgment conditions, thereby obtaining frequency domain coefficients; the frequency domain transformation method includes the Hadamard transform method based on addition and subtraction operations; An energy value determination unit is used to calculate the sum of the absolute values of the frequency domain coefficients and use the sum of the absolute values of the frequency domain coefficients as an energy value characterizing the texture complexity of the image region to be encoded; the larger the energy value, the more complex the texture of the image region to be encoded. The data acquisition unit is used to multiply the energy value and a preset weight coefficient to obtain a texture masking component; the texture masking component is proportional to the energy value.
8. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor for executing the computer program to implement the HDR video coding method based on texture-luminance-aware masking as described in any one of claims 1 to 5.
9. A computer-readable storage medium, characterized in that, Used to store a computer program, which, when executed by a processor, implements the HDR video coding method based on texture-luminance-aware masking as described in any one of claims 1 to 5.
Citation Information
Patent Citations
CN103124347A
CN117241027A