Self-adaptive real-time image quality enhancement method and system

Through the GPU acceleration processing framework and adaptive image quality enhancement technology, the high cost and delay problems in video quality enhancement are solved, and low-cost, real-time and efficient video quality improvement is achieved, especially in skin tone protection and adaptive adjustment.

CN120281915APending Publication Date: 2025-07-08SHENZHEN MAPLE LEAF INTERACTIVE TECHNOLOGY CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510464308.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-14
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

The prior art has problems such as high cost, delay, inflexible fixed parameter processing, difficulty in balancing real-time and complexity, and insufficient skin tone protection in video quality, affecting user experience and video quality.

Method used

The GPU acceleration processing framework is adopted, combining image data conversion, adaptive brightness enhancement, multi-frame combined skin color detection, local adaptive color enhancement, dimensionality reduction bilateral filtering and texture adaptive detail enhancement modules, and adaptive image quality enhancement is achieved through OpenGL and CNN models.

Benefits of technology

It achieves low-cost, real-time and efficient video quality improvement, adaptive adjustments avoid overexposure and noise amplification, protects the skin color area, and improves the naturalness and overall quality of the video.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120281915A_ABST
    Figure CN120281915A_ABST
Patent Text Reader

Abstract

The invention provides a self-adaptive real-time image quality enhancement method and system, and the system comprises an image data conversion module which is used for converting YUV format data of a video stream into an RGB format through OpenGL, and further converting the RGB format into an HSL color space; the adaptive brightness enhancement module dynamically adjusts an enhancement coefficient according to the brightness component of the current frame; the multi-frame joint skin color detection module is used for determining a skin color area based on the chromaticity and saturation threshold range of the HSL color space; the local self-adaptive color enhancement module is used for inputting image generation dynamic parameters and performing color enhancement; the noise analysis and noise reduction module is used for performing soft threshold processing on the high-frequency signal; the texture adaptive detail enhancement module is used for adjusting sharpening strength based on texture complexity; and the GPU acceleration processing framework is used for rendering and outputting the enhanced RGB image in real time. Through a series of technical innovations, the image quality of video playing and the user experience are remarkably improved in multiple aspects, and the method has important practical application value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of information image processing, and particularly relates to a method and system for adaptive real-time image quality enhancement. Background Art

[0002] With the popularization of the Internet and the popularity of short videos, video content has become increasingly diverse. Many Internet companies are targeting the overseas short drama market, including regions such as Europe, America, South America, Southeast Asia, the Middle East, Japan, and South Korea. However, due to the diverse video sources, the video quality varies greatly. Common problems include too dark brightness, too low saturation, and lack of picture details, which seriously affect the user's viewing experience.

[0003] Deficiencies of the prior art: 1. Limitations of server-side processing: High cost: Performing image quality enhancement processing on the server requires additional computing resources, increasing the server load and operating costs.

[0004] Latency issue: Server-side processing takes time and cannot immediately respond to the demand for real-time playback, affecting the user experience.

[0005] 2. Problem of fixed parameters: Lack of flexibility: Traditional solutions use fixed parameters to uniformly process the entire image and cannot adaptively adjust according to different picture features (such as brightness, saturation, noise, etc.), resulting in overexposure, noise amplification, or color imbalance in some areas.

[0006] Poor visual effect: Especially in the face area, if uniform color enhancement parameters are used, it may lead to problems such as yellowish skin color and unnatural appearance.

[0007] 3. Balance between real-time performance and complexity: Difficulty in real-time processing: To achieve real-time processing, traditional noise reduction methods such as mean filtering and median filtering are simple but result in loss of details, while BM3D (Block Matching and 3D Filtering) and NLM (Non-Local Means) have good filtering effects but high computational complexity and are difficult to achieve real-time processing.

[0008] Insufficient local optimization: Most existing methods uniformly process the entire image and lack targeted optimization for local areas (such as areas with complex textures or large noise).

[0009] 4. Challenges in skin color protection: Problems in color enhancement: When performing global color enhancement, if skin tone protection is not considered, it may cause color changes in the skin tone area, making the face look unnatural and affecting the visual experience.

[0010] Therefore, the existing technology has deficiencies and needs further improvement. Summary of the Invention

[0011] In view of the problems existing in the prior art, the present invention provides a method and system for adaptive real-time image quality enhancement.

[0012] To achieve the above object, the specific solution of the present invention is as follows: The present invention provides a system for adaptive real-time image quality enhancement, including: A GPU acceleration processing framework for receiving YUV format data of a client video stream and performing full-process parallel computing on the GPU side; An image data conversion module that converts the YUV format to the RGB format through OpenGL and further converts the RGB format to the HSL color space, separating the luminance (L), saturation (S), and chrominance (H) components; An adaptive luminance enhancement module that dynamically calculates a non-linear enhancement coefficient based on the luminance component of the current frame and performs piecewise limiting processing on the luminance value, where the enhancement coefficient in the area with luminance lower than 0.8 is inversely proportional to the luminance, and the enhanced luminance value is limited to ≤1; A multi-frame joint skin color detection module that initially delimits the skin color area in the HSL color space based on the threshold ranges of chrominance H∈[0°, 36°] and saturation S∈[0.2, 0.6], and optimizes the skin color area mask by fusing the detection results of the adjacent 3 previous frames and 3 subsequent frames; A local adaptive color enhancement module, including a pre-trained CNN model and a traditional saturation adjustment unit, where the CNN model is trained with low-saturation - high-saturation image pairs, outputs a scene-adaptive saturation gain coefficient, and enhances only the non-skin color area; The CNN model includes: Lightweight FCN, an attention mechanism U-shaped network, and a parameter regression model; A dimensionality reduction bilateral filtering module that performs wavelet decomposition on the image, performs one-dimensional bilateral filtering on the low-frequency signal along the horizontal or vertical direction, and uses soft threshold denoising for the high-frequency signal, where the kernel function weights of the one-dimensional filtering are jointly determined by the spatial distance and the pixel intensity difference; A texture adaptive detail enhancement module that divides the image into high, medium, and low texture complexity regions based on the gradient magnitude and dynamically adjusts the sharpening intensity, where the sharpening intensity is positively correlated with the gradient magnitude; A real-time rendering module that inverse-converts the processed HSL data to the RGB format and renders and outputs it through the GPU shader pipeline.

[0013] Further, the calculation of the non-linear enhancement coefficient satisfies the following piecewise function:

[0014] where L: the luminance component of the current pixel, with a value range of [0, 1]; α(L): the luminance enhancement coefficient, which is inversely proportional to the luminance value; The enhanced luminance value is L′ = min(1.0, L + α(L)).

[0015] Further, the kernel function weight of the one-dimensional bilateral filter is:

[0016] where x, y: the coordinates of the pixels within the filtering window; : the spatial distance between pixel coordinates; I(x), I(y): the intensity values of pixels x and y; σd: the spatial distance weight parameter, with a value of 1.5; σr: the pixel intensity difference weight parameter, with a value of 10.0; σd = 1.5 controls the spatial distance weight, σr = 10.0 controls the pixel intensity difference weight, and the filtering direction is only executed unidirectionally along the horizontal or vertical direction.

[0017] Further, the training data of the CNN model includes the following paired samples: Input: a low-saturation image (S ≤ 0.3) and its corresponding scene labels, including night scene, backlight, and indoor; Output: a saturation gain coefficient matrix, where each element in the matrix corresponds to the gain value of an image block, and the gain coefficient in the non-skin-color area is 1.5 - 2.0.

[0018] Further, the division of the texture complexity region satisfies: High complexity region: gradient magnitude ≥ 60, sharpening intensity amount = 1.5; Gradient magnitude: the image gradient magnitude calculated by the Sobel operator, which is used to measure the texture complexity; amount: the sharpening intensity coefficient, with a value range of 0 - 1.5; Medium complexity region: 30 ≤ gradient magnitude < 60, amount = 1.0; Low complexity region: gradient magnitude < 30, amount = 0.

[0019] The present invention also provides a method for adaptive real-time image quality enhancement. Based on the above system, the method includes the following steps: Step S1: Convert the YUV format data of the client video stream to the RGB format on the GPU side, and further convert it to the HSL color space; Step S2: Calculate the non-linear enhancement coefficient α(L) according to the luminance component L of the current frame, enhance the luminance in the region where L ≤ 0.8, and limit the enhanced luminance value to L' ≤ 1; Step S3: Detect the skin color region with chromaticity H ∈ [0°, 36°] and saturation S ∈ [0.2, 0.6] in the HSL space, and generate an optimized skin color mask by combining the detection results of the adjacent 3 frames before and 3 frames after; Step S4: Generate a saturation gain coefficient matrix for non-skin color regions through a pre-trained CNN model, and apply a traditional saturation adjustment algorithm for local enhancement; Step S5: Divide the image into several blocks, calculate the variance σ² of each block, perform wavelet decomposition on the blocks with σ² ≥ 100, perform one-dimensional bilateral filtering on the low-frequency signal along the horizontal or vertical direction, and perform soft threshold denoising on the high-frequency signal; Step S6: Calculate the image gradient magnitude, divide the high, medium, and low texture complexity regions, and dynamically adjust the sharpening intensity amount according to the gradient magnitude; Step S7: Inversely convert the processed HSL data to the RGB format, and perform real-time rendering output through the GPU shader pipeline, with a frame rate ≥ 30fps.

[0020] Further, in the step S3, the optimization of the skin color mask is calculated by the following formula:

[0021] Among them, Mt: The skin color region mask preliminarily detected in the t-th frame, with a value of 0 for non-skin color or 1 for skin color; Mt−3, Mt−2, …, Mt+3: The skin color masks of the adjacent 3 frames before and 3 frames after; Denominator 7: The total number of frames participating in the joint optimization, the current frame + 3 frames before and after; Mt is the skin color region preliminarily detected in the t-th frame. When the optimized mask value ≥ 0.6, it is determined as the skin color region.

[0022] Further, in the step S5, the formula for soft threshold processing is:

[0023] Among them, x: The wavelet coefficient of the high-frequency signal; T: The dynamic threshold, which is positively correlated with the noise intensity. When the noise intensity ≥ 60, T = 15; sign(x): Take the sign of x, +1 or -1.

[0024] Furthermore, in the step S6, the dynamic adjustment of the sharpening intensity satisfies: amount = 0.025 × gradient_magnitude And the maximum value of amount is limited to 1.5; gradient_magnitude: The image gradient magnitude, calculated by the Sobel operator.

[0025] Furthermore, in the steps S1 to S7, the parallel computing on the GPU side is achieved through the following optimizations: The YUV→RGB conversion uses the fragment shader of OpenGL ES 3.0; The one-dimensional bilateral filtering is processed in parallel by blocks through the GPU compute shader Compute Shader; The rendering output adopts a double-buffer mechanism, with a latency ≤ 30 ms.

[0026] Adopting the technical solution of the present invention has the following beneficial effects: 1. Improve the viewing experience Real-time processing: By enhancing the picture quality during client playback, the video quality can be improved in real time without waiting for the server to process, greatly improving the user's viewing experience.

[0027] Adaptive adjustment: According to different picture features (such as brightness, saturation, noise, etc.), it makes adaptive adjustments, avoiding problems such as overexposure, noise amplification, or color imbalance caused by fixed parameter processing.

[0028] 2. Significant cost-effectiveness Reduce the server load: Transfer the work of picture quality enhancement from the server to the client, reducing the consumption of the server's computing resources and operating costs.

[0029] Optimize resource allocation: Using the GPU for image processing is more efficient than CPU processing, further reducing energy consumption and costs.

[0030] 3. Skin color protection and natural look Skin color detection and protection: Detect the skin color through the HSL color space and protect the detected skin color area to prevent color deviation during color enhancement, ensuring that the face area looks natural and real.

[0031] Multi-frame joint optimization: Adopt multi-frame joint optimization technology to improve the accuracy of skin color detection and enhance the naturalness of the overall visual effect.

[0032] 4. Local Adaptive Processing Local color enhancement: By combining the CNN convolutional neural network with traditional methods, local adaptive color enhancement of non-skin-color regions is achieved, improving the saturation and color richness of the image.

[0033] Local noise reduction and detail enhancement: By performing noise estimation and local noise reduction on specific regions, and enhancing details according to texture complexity, both the details of the image are retained and the noise is effectively removed, improving the overall quality of the image.

[0034] 5. Efficient Real-time Processing Fast noise reduction: By using bilateral filtering and dimensionality reduction processing, efficient real-time noise reduction is achieved while retaining the details of the image, meeting the requirements of real-time processing.

[0035] Optimized filtering algorithm: By wavelet decomposition and soft threshold processing, low-frequency and high-frequency signals are separated, improving the filtering speed and further enhancing the processing efficiency. 6. Comprehensive Image Quality Enhancement Multi-dimensional enhancement: Covering multiple aspects such as brightness enhancement, color enhancement, noise reduction processing, and detail enhancement, the quality of the video is comprehensively improved.

[0036] Flexible application: Applicable to video content under various different lighting conditions, whether it is a low-light scene or a high-brightness scene, it can provide high-quality visual effects.

[0037] 7. Technical Innovation and Practicality Innovative design: The present invention proposes a series of innovative design solutions, including adaptive brightness enhancement, local color enhancement, skin-color protection, noise estimation and reduction, detail enhancement, etc. These designs not only solve the deficiencies in the prior art but also provide highly practical technical solutions.

[0038] Easy to integrate: The system and method can be easily integrated into existing video players without large-scale modification of the existing architecture, facilitating promotion and application. Description of the Drawings

[0039] Figure 1 is the principle block diagram of the present invention; Figure 2 is the overall flowchart of the present invention. Detailed Description of the Preferred Embodiment

[0040] The present invention will be further described in detail below with reference to the drawings and embodiments; it can be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention; in addition, it should be noted that for the sake of description, only parts related to the present invention are shown in the drawings rather than all of them.

[0041] Combined Figure 1 - Figure 2 As shown, the present invention provides a system for adaptive real-time picture quality enhancement, including: A GPU acceleration processing framework for receiving YUV format data of a client video stream and performing full-process parallel computing on the GPU side; An image data conversion module that converts the YUV format to the RGB format through OpenGL and further converts the RGB format to the HSL color space, separating the luminance (L), saturation (S), and chrominance (H) components; An adaptive luminance enhancement module that dynamically calculates a non-linear enhancement coefficient according to the luminance component of the current frame and performs piecewise clipping processing on the luminance value, where the enhancement coefficient in the area with luminance lower than 0.8 is inversely proportional to the luminance, and the enhanced luminance value is limited to ≤1; A multi-frame combined skin color detection module that preliminarily delimits the skin color area based on the threshold range of chrominance H∈[0°, 36°] and saturation S∈[0.2, 0.6] in the HSL color space, and fuses the detection results of the adjacent 3 previous frames and 3 subsequent frames to optimize the skin color area mask; A local adaptive color enhancement module including a pre-trained CNN model and a traditional saturation adjustment unit. The CNN model is trained by low-saturation - high-saturation image pairs, outputs a scene-adaptive saturation gain coefficient, and enhances only non-skin color areas; The CNN model includes: Lightweight FCN, an attention mechanism U-shaped network, and a parameter regression model; A dimensionality reduction bilateral filtering module that performs wavelet decomposition on the image, performs one-dimensional bilateral filtering on the low-frequency signal along the horizontal or vertical direction, and uses soft threshold denoising for the high-frequency signal, where the kernel function weights of the one-dimensional filtering are jointly determined by the spatial distance and the pixel intensity difference; A texture adaptive detail enhancement module that divides the image into high, medium, and low texture complexity regions based on the gradient magnitude and dynamically adjusts the sharpening intensity, where the sharpening intensity is positively correlated with the gradient magnitude; A real-time rendering module that inversely converts the processed HSL data to the RGB format and renders and outputs it through the GPU shader pipeline.

[0042] The calculation of the non-linear enhancement coefficient satisfies the following piecewise function:

[0043] Where, L: The luminance component of the current pixel, with a value range of [0, 1]; α(L): The luminance enhancement coefficient, which is inversely proportional to the luminance value; The enhanced luminance value is L′ = min(1.0, L + α(L)).

[0044] The kernel function weights of the one-dimensional bilateral filtering are as follows:

[0045] where x, y: the coordinates of pixels within the filtering window; : the spatial distance between pixel coordinates; I(x), I(y): the intensity values of pixels x and y; σd: the spatial distance weight parameter, with a value of 1.5; σr: the pixel intensity difference weight parameter, with a value of 10.0; σd = 1.5 controls the spatial distance weight, σr = 10.0 controls the pixel intensity difference weight, and the filtering direction is only executed unidirectionally along the horizontal or vertical direction.

[0046] The training data of the CNN model includes the following paired samples: Input: low-saturation images (S ≤ 0.3) and their corresponding scene labels, including night scene, backlight, and indoor; Output: a saturation gain coefficient matrix, where each element in the matrix corresponds to the gain value of an image block, and the gain coefficient for non-skin regions is 1.5 - 2.0.

[0047] The division of the texture complexity regions satisfies the following: High complexity region: gradient magnitude ≥ 60, sharpening intensity amount = 1.5; Gradient magnitude: the image gradient magnitude calculated by the Sobel operator, used to measure texture complexity; amount: the sharpening intensity coefficient, with a value range of 0 - 1.5; Medium complexity region: 30 ≤ gradient magnitude < 60, amount = 1.0; Low complexity region: gradient magnitude < 30, amount = 0.

[0048] The present invention also provides a method for adaptive real-time image quality enhancement. Based on the above system, the method includes the following steps: Step S1: Convert the YUV format data of the client video stream to the RGB format on the GPU side and further convert it to the HSL color space; Step S2: Calculate the non-linear enhancement coefficient α(L) according to the current frame luminance component L, perform luminance enhancement on the region where L ≤ 0.8, and limit the enhanced luminance value to L' ≤ 1; Step S3: Detect the skin-color regions with hue H ∈ [0°, 36°] and saturation S ∈ [0.2, 0.6] in the HSL color space, and generate an optimized skin-color mask by combining the detection results of the adjacent 3 frames before and 3 frames after; Step S4: Generate a saturation gain coefficient matrix for non-skin-color regions through a pre-trained CNN model, and apply a traditional saturation adjustment algorithm for local enhancement; Step S5: Divide the image into several blocks, calculate the variance σ² of each block, perform wavelet decomposition on the blocks with σ² ≥ 100, perform one-dimensional bilateral filtering on the low-frequency signals along the horizontal or vertical direction, and perform soft-threshold denoising on the high-frequency signals; Step S6: Calculate the gradient magnitude of the image, divide the regions with high, medium, and low texture complexities, and dynamically adjust the sharpening intensity amount according to the gradient magnitude; Step S7: Inversely convert the processed HSL data into the RGB format, and perform real-time rendering output through the GPU shader pipeline with a frame rate ≥ 30fps.

[0049] In the said Step S3, the optimization of the skin-color mask is calculated by the following formula:

[0050] Wherein, Mt: The skin-color region mask initially detected in the t-th frame, with a value of 0 for non-skin-color or 1 for skin-color; Mt−3, Mt−2, …, Mt+3: The skin-color masks of the adjacent 3 frames before and 3 frames after; The denominator 7: The total number of frames participating in the joint optimization, the current frame + 3 frames before and after; Mt is the skin-color region initially detected in the t-th frame. When the optimized mask value ≥ 0.6, it is determined as a skin-color region.

[0051] In the said Step S5, the formula for the soft-threshold processing is:

[0052] Wherein, x: The wavelet coefficient of the high-frequency signal; T: The dynamic threshold, which is positively correlated with the noise intensity. When the noise intensity ≥ 60, T = 15; sign(x): Take the sign of x, +1 or -1.

[0053] In the said Step S6, the dynamic adjustment of the sharpening intensity satisfies: amount = 0.025 × gradient_magnitude And the maximum value of amount is limited to 1.5; gradient_magnitude: The magnitude of the image gradient, calculated using the Sobel operator; amount: The sharpening intensity coefficient, with a maximum limit of 1.5.

[0054] In the steps S1 to S7, the parallel computing on the GPU side is achieved through the following optimizations: The YUV→RGB conversion uses the fragment shader of OpenGL ES 3.0; The one-dimensional bilateral filtering is processed in parallel by blocks using the GPU compute shader Compute Shader; The rendering output adopts a double-buffer mechanism with a latency ≤ 30 ms.

[0055] Specifically as follows: 1. The CPU side uploads the YUV (a format describing the image color encoding, where Y represents luminance and UV represents chrominance) image data to the GPU side. Based on OpenGL (Open Graphics Library, an open-source graphics library), the YUV is converted into the RGB (a format for image rendering) color space, as shown in Equation 1: YUV to RGB matrix: R = 1.164 * Y + 1.792 * V G = 1.164 * Y - 0.213 * U - 0.533 * V B = 1.164 * Y + 2.112 * U 2. Then the RGB is converted into the HSL (where H represents hue, S represents saturation, and L represents luminance) color space, as shown in Equation 2, to obtain the saturation component S and the luminance component L.

[0056] Calculate the maximum and minimum values:

[0057]

[0058] Calculate the hue H:

[0059] Calculate the saturation:

[0060] Calculate the luminance L:

[0061] 3. During video acquisition and shooting, due to too dim light or backlit shooting, the picture is dark, and there are differences in brightness in different regions. According to the current brightness adaptive adjustment coefficient, the smaller the brightness, the larger the coefficient, and a clipping process is performed to avoid exceeding the maximum value of 1. When the brightness is 0 (black picture), no processing is done. When the brightness is greater than 0.8, it is considered that the brightness is sufficient and no processing is done to avoid overexposure.

[0062] 4. Perform skin color detection on the current image frame. Since the RGB color space is easily affected by the light intensity, while the HSL color space performs well under different lighting conditions, it is converted to the HSL color space for skin color detection. The color value range of the skin color is shown in Formula 3. After matching the skin color range, the skin color region is initially obtained. Then, multi-frame joint optimization is adopted, that is, combining the detection results of adjacent front and rear frame images to improve the accuracy, and finally determine whether it is the skin color region.

[0063] Formula 3: Skin color detection range based on HSL

[0064]

[0065]

[0066] 5. During video acquisition and shooting, due to too dim light, the colors of the picture are dull, and the overall quality of the video is greatly reduced. To improve the picture saturation and color richness, color enhancement is proposed here. Skin color protection is performed for the skin color region. If it falls within the skin color region, no color enhancement is done to avoid color deviation in the skin color region and yellowing of the face, which appears unnatural and untrue.

[0067] There are two methods for color enhancement. The first is the traditional method of adjusting the saturation. In this way, fixed parameters are used for adjustment, which is applied to the entire video frame, resulting in possible over-saturation and color mutation in some frames. The second is the method based on neural network, using an AI (Artificial Intelligence) model for color enhancement, but it takes a relatively long time and is not suitable for real-time color enhancement scenarios.

[0068] Here, a combination of CNN convolutional neural network and the traditional method is proposed. First, based on the method of CNN convolutional neural network, two image sets of dull color and bright color are used for training to obtain empirical parameters for different frames and different scenarios. Then, the empirical parameters obtained from the color enhancement model are applied to the traditional method for color enhancement, so as to achieve local adaptive color enhancement and at the same time achieve the effect of real-time processing.

[0069] 6. Based on the HSL color space, after performing brightness enhancement and color enhancement, reverse the HSL color space to the RGB color space, as shown in Equation 4: HSL to RGB Calculate intermediate variables:

[0070]

[0071]

[0072] Calculate RGB process variables:

[0073] Calculate the final RGB value:

[0074] 7. According to the analysis of a large number of images, in low-light scenes, the noise is greater than that in normal brightness; in flat areas, the noise is more obvious than in non-flat areas. Since noise estimation and image denoising are time-consuming, for the purpose of real-time processing, first perform scene analysis to initially delimit the areas with high noise. Then perform noise estimation on specific areas, divide the image into multiple 16x16 blocks, and calculate their mean, sum of squares, and variance, as shown in Equation 5. Use the variance to estimate the noise, and map the noise result to [0, 100]. If the noise in this area is relatively large, perform local denoising.

[0075] Equation 5, Noise Estimation:

[0076] Traditional image denoising methods include mean filtering, median filtering, BM3D filtering (Block matching and 3D filtering), and NLM filtering (Non-Local Means). Mean filtering and median filtering are relatively simple, but the disadvantage is that details are lost, resulting in a decrease in image quality. The effects of BM3D and NLM filtering are relatively good, but the disadvantages are high complexity and difficulty in achieving real-time processing. Based on the trade-off between denoising effect and complexity, bilateral filtering is selected here, which can not only retain details to achieve a good denoising effect but also perform real-time processing. First, perform wavelet decomposition to separate the low-frequency signal and the high-frequency signal. The high-frequency signal is subjected to soft threshold processing, and the low-frequency signal is denoised using bilateral filtering. Traditional bilateral filtering is performed in a two-dimensional space and is time-consuming. Here, dimensionality reduction processing is used to reduce the two-dimensional space to a one-dimensional space, and the filtering is optimized in the directions from top to bottom, from bottom to top, from left to right, and from right to left, thereby improving the filtering speed.

[0077]

[0078] ​8. Analyze the texture structure of the image to obtain the complexity of the texture region. Regions with low texture complexity are approximately flat areas and do not require detail enhancement, saving computational time. For regions with high texture complexity, local detail enhancement is performed. Adjust the sharpening intensity amount according to the texture complexity to achieve local adaptive detail enhancement.

[0079] The process of detail enhancement is shown in Equation 6. First, perform Gaussian filtering to obtain the blur filtered image, then subtract the blur filtered image from the origin original image to get the difference. Next, multiply the difference by the sharpening intensity amount, and finally add it to the original image to obtain the sharpened image after detail enhancement.

[0080] Equation 6, Detail enhancement based on unsharp mask: sharpen = origin + amount * (origin - blur) 9. After completing all image quality enhancements, render the RGB image to finally achieve real-time adaptive image quality enhancement.

[0081] The working principle of the present invention: 1. Overall process framework The present invention is based on real-time processing on the client GPU side. Through a multi-stage adaptive enhancement algorithm, the image quality is optimized frame by frame during video playback. The core process is divided into the following stages: 1.1 Data preprocessing and color space conversion 1.2 Brightness and color enhancement 1.3 Noise suppression and detail restoration 1.4 GPU accelerated rendering output 2. Working principle of core module collaboration 2.1 Data preprocessing and color space conversion Input data: The client video stream is in YUV format (luminance Y + chrominance UV components).

[0082] YUV → RGB conversion: Execute the matrix operation of Equation 1 on the GPU side through OpenGL to convert YUV to RGB format, providing basic data for subsequent rendering and enhancement.

[0083]

[0084] RGB → HSL conversion: Further convert RGB to the HSL color space, separating the luminance (L), saturation (S), and chrominance (H) components (Equation 2) for easy partition processing.

[0085] 2.2 Adaptive Brightness Enhancement Dynamic adjustment strategy: Calculate the enhancement coefficient based on the brightness component of the current frame (L ∈ [0, 1]): When L = 0 (pure black) or L > 0.8, no processing is performed; When 0 < L ≤ 0.8, the enhancement coefficient is inversely proportional to the brightness (the smaller L, the larger the coefficient).

[0086] Clipping processing: Limit the enhanced brightness value to ≤ 1 to avoid overexposure.

[0087] Local optimization: For images with uneven brightness distribution (such as backlight scenes), only enhance the dark areas, and keep the bright areas at their original values.

[0088] 2.3 Multi-frame Joint Skin Color Detection and Protection Skin color detection logic: HSL threshold screening: Define the skin color area in the HSL space (H ∈ [0°, 36°], S ∈ [0.2, 0.6]).

[0089] Multi-frame joint optimization: Combine the detection results of adjacent front and back frames to eliminate single-frame misdetection (such as color deviation caused by sudden changes in light).

[0090] Protection mechanism: Disable saturation enhancement for skin color areas to prevent the face from turning yellow (ΔE ≤ 2.0).

[0091] 2.4 Local Adaptive Color Enhancement Hybrid enhancement architecture: CNN model pre-training: Use low-saturation - high-saturation image pairs to train the convolutional neural network to generate scene-based enhancement parameters.

[0092] Dynamic parameter application: Combine the parameters output by the CNN (such as the saturation gain coefficient) with traditional algorithms, and only enhance non-skin color areas.

[0093] Effect example: Input saturation S = 0.3 → Enhanced S = 0.6 (ideal visual range 0.5 - 0.7).

[0094] 2.5 Noise Assessment and Dimensionality Reduction Noise Reduction Noise intensity assessment: Divide the image into 16×16 blocks and calculate the variance of each block (Formula 5):

[0095] (μ: mean, N = 256) Map the noise intensity to a score of [0, 100], and trigger noise reduction when the score is higher than 40 points.

[0096] Dimensionality Reduction Bilateral Filtering: Wavelet Decomposition: Separate low-frequency (smooth area) and high-frequency (detail / noise) signals.

[0097] Low-frequency Processing: Perform one-dimensional bilateral filtering on the low-frequency signal along the horizontal / vertical direction, reducing the computational complexity by 60% compared to two-dimensional.

[0098] High-frequency Processing: Apply a soft threshold to the high-frequency signal (formula: , to suppress noise.

[0099] 2.6 Texture Adaptive Detail Enhancement Dynamic Adjustment of Sharpening Intensity: Calculate the gradient magnitude of the image and divide the texture complexity levels (low / medium / high).

[0100] Apply strong sharpening (amount = 1.5) to high-complexity areas (such as hair, building edges), and do not process flat areas (such as the sky) (amount = 0).

[0101] Sharpening Formula: Sharpened Image = Original Image + amount × (Original Image - Gaussian Blurred Image) 2.7 GPU Full Process Acceleration Parallel Design: Assign computationally intensive tasks such as color space conversion, filtering, and CNN inference to multi-core parallel processing on the GPU.

[0102] Transfer Optimization: Asynchronous transfer of YUV data through the dual PBO of the GPU to reduce the data transfer latency from the CPU to the GPU Real-time Guarantee: The processing time for a single frame ≤ 15ms (1080p resolution), and the frame rate ≥ 30fps.

[0103] 3. Key Technical Innovation Principles

[0104] 4. Schematic Diagram of the Working Principle: YUV Input → GPU Convert to RGB → Convert to HSL → [Brightness Enhancement → Skin Color Detection → Color Enhancement] → Noise Reduction → Detail Enhancement → Inverse Convert RGB → Render Output The arrow indicates the data flow direction, and the content in the square brackets is the parallel processing branch.

[0105] Example 1: Real-time Enhancement of Low-brightness Videos Application Scenario: Short videos taken by users at night (resolution 1080p, brightness L = 0.2, noise intensity 65).

[0106] Processing Flow: YUV→RGB Conversion: Converted by OpenGL on the GPU side according to Formula 1, taking 2 ms.

[0107] RGB→HSL Conversion: Extract the luminance component L = 0.2, triggering the adaptive brightness enhancement module.

[0108] Brightness Enhancement: Calculate the enhancement coefficient 1.5 (L = 0.2), and the output luminance after clipping is L = 0.45.

[0109] Noise Reduction Processing: Divide into 16×16 blocks to detect noise (variance σ² = 120), perform one-dimensional bilateral filtering on low-frequency signals, and soft-threshold processing on high-frequency signals, reducing the noise intensity to 30.

[0110] Rendering Output: The total processing time is 14 ms, and the frame rate is stable at 30 fps.

[0111] Effect Comparison: The PSNR is increased from 28 dB to 33 dB, and the visibility of dark details is increased by 40%.

[0112] User surveys show that the satisfaction level has increased from 2.8 / 5.0 to 4.3 / 5.0. Example 2: Portrait Video Processing with Skin Tone Protection Application Scenario: Portrait video taken outdoors against the backlight (face area H = 20°, S = 0.5, background saturation S = 0.3).

[0114] Processing Flow: Skin Tone Detection: Define the skin tone area as H ∈ [15°, 25°], S ∈ [0.4, 0.6] in the HSL space.

[0115] Multi-frame Optimization: Combine the detection results of the previous and next 3 frames to correct single-frame misdetections (such as sudden changes in H value caused by light reflection).

[0116] Color Enhancement: Non-skin Tone Areas: Apply the saturation gain coefficient of 1.8 generated by CNN, increasing S from 0.3 to 0.54; Skin Tone Areas: Keep S = 0.5, with ΔE color difference ≤ 1.5.

[0117] Output Effect: The face looks natural without yellowing, the background color is vivid, and the processing delay is 18 ms.

[0118] Data Support: The skin tone detection accuracy is 95%, and the misdetection rate < 3%.

[0119] The SSIM (structural similarity) of the background area is increased from 0.75 to 0.92. Example 3: Restoration of Old Movies with High Noise Application Scenario: Film movies in the 1980s (noise intensity 70, resolution 720p).

[0121] Processing Flow: Noise Assessment: Divide into 16×16 blocks and calculate the variance σ² = 200 (noise intensity 70).

[0122] Dimensionality-Reduced Bilateral Filtering: Low-Frequency Signal: Filter in the horizontal direction, taking 8 ms (traditional two-dimensional filtering takes 20 ms); High-Frequency Signal: Soft threshold T = 15 to suppress granular noise.

[0123] Detail Enhancement: Apply sharpening with amount = 1.2 to the building area with complex texture (gradient magnitude > 50).

[0124] Effect Comparison: The noise intensity is reduced to 25, and the PSNR is increased by 6 dB; The edge detail retention rate is 90%, while the traditional median filtering is only 65%.

[0125] Example 4: Enhancement of Natural Scenery with Complex Textures Application Scenario: 4K natural scenery video (high leaf texture complexity, original sharpening intensity fixed at 1.0).

[0126] Processing Flow: 1. Texture Analysis: Calculate the gradient magnitude and divide the texture levels: High-Complexity Region (leaves): Gradient > 60, amount = 1.5; Low-Complexity Region (sky): Gradient < 10, amount = 0.

[0127] 2. Sharpening Processing: Leaf Area = Original Image + 1.5×(Original Image - Gaussian Blurred Image) 3. GPU Rendering: The total time is 22 ms (4K resolution), and the frame rate is stable at 30 fps.

[0128] Effect Data: The SSIM of the texture area is increased by 25%, and the artifact rate is reduced to less than 2%.

[0129] Example 5: Verification of the Full-Process Efficiency of End-Side GPU Test Environment: Mid-range mobile phone (Snapdragon 778G, GPU Adreno 642L).

[0130] Test Content: Process a 30fps 1080p video stream and compare it with the traditional CPU solution.

[0131] Processing flow: Full GPU process: YUV→RGB conversion, HSL processing, noise reduction, and the entire rendering process are GPU-parallelized with a utilization rate of 92%; The time consumption per frame is 15 ms, and the memory occupancy is 180 MB.

[0132] CPU comparison group: The same algorithm is implemented with OpenCV-CPU, with a time consumption of 50 ms per frame and a memory occupancy of 450 MB.

[0133] Performance data: GPU frame rate: 30 fps (smooth), CPU frame rate: ≤20 fps (obvious lag); Power consumption: The GPU solution is 3.2 W, and the CPU solution is 5.8 W (a 45% reduction). The above are only the preferred embodiments of the present invention, and do not limit the scope of the present invention. Any equivalent structural transformation made under the inventive concept of the present invention, or direct / indirect application in other related technical fields, is included in the protection scope of the present invention.

Claims

1. An adaptive real-time picture quality enhancement system, characterized in that, Including: A GPU-accelerated processing framework for receiving YUV format data of a client video stream and performing full-process parallel computing on the GPU side; An image data conversion module that converts the YUV format to the RGB format through OpenGL and further converts the RGB format to the HSL color space, separating the luminance L, saturation S, and chrominance H components; An adaptive brightness enhancement module that dynamically calculates a non-linear enhancement coefficient based on the current frame's luminance component and performs piecewise clipping processing on the luminance value, where the enhancement coefficient in the area with luminance below 0.8 is inversely proportional to the luminance, and the enhanced luminance value is limited to ≤1; A multi-frame combined skin color detection module that initially delimits the skin color area based on the threshold ranges of chrominance H ∈ [0°, 36°] and saturation S ∈ [0.2, 0.6] in the HSL color space and optimizes the skin color area mask by fusing the detection results of the adjacent 3 previous frames and 3 subsequent frames; A local adaptive color enhancement module including a pre-trained CNN model and a traditional saturation adjustment unit. The CNN model is trained with low-saturation - high-saturation image pairs, outputs a scene-adaptive saturation gain coefficient, and enhances only non-skin color areas; The CNN model includes: Lightweight FCN, an attention mechanism U-shaped network, and a parameter regression model; A dimensionality reduction bilateral filtering module that performs wavelet decomposition on the image, performs one-dimensional bilateral filtering on the low-frequency signal along the horizontal or vertical direction, and uses soft threshold denoising for the high-frequency signal, where the kernel function weights of the one-dimensional filtering are jointly determined by the spatial distance and the pixel intensity difference; A texture adaptive detail enhancement module that divides the image into high, medium, and low texture complexity regions based on the gradient magnitude and dynamically adjusts the sharpening intensity, where the sharpening intensity is positively correlated with the gradient magnitude; A real-time rendering module that inversely converts the processed HSL data to the RGB format and renders and outputs it through the GPU shader pipeline.

2. The system according to claim 1, wherein: The calculation of the non-linear enhancement coefficient satisfies the following piecewise function: Wherein, L: The luminance component of the current pixel, with a value range of [0, 1]; α(L): The luminance enhancement coefficient, which is inversely proportional to the luminance value; The enhanced luminance value is L′ = min(1.0, L + α(L)).

3. The system according to claim 1, wherein: The kernel function weights of the one-dimensional bilateral filtering are: Wherein, x, y: The coordinates of the pixels within the filtering window; ∥x - y∥: The spatial distance between the pixel coordinates; I(x), I(y): The intensity values of pixels x and y; σd: The spatial distance weight parameter, with a value of 1.5; σr: The pixel intensity difference weight parameter, with a value of 10.0; σd = 1.5 controls the spatial distance weight, σr = 10.0 controls the pixel intensity difference weight, and the filtering direction is only performed unidirectionally along the horizontal or vertical direction.

4. The system according to claim 1, wherein: The training data of the CNN model includes the following paired samples: Input: Low-saturation images (S ≤ 0.3) and their corresponding scene labels, including night scene, backlight, and indoor; Output: Saturation gain coefficient matrix, where each element in the matrix corresponds to the gain value of an image block, and the gain coefficient of non-skin-color regions is 1.5 - 2.

0.

5. The system according to claim 1, characterized in that: The division of the texture complexity regions satisfies: High complexity region: Gradient magnitude ≥ 60, sharpening intensity amount = 1.5; Gradient magnitude: The image gradient magnitude calculated by the Sobel operator, used to measure texture complexity; amount: Sharpening intensity coefficient, with a value range of 0 - 1.5; Medium complexity region: 30 ≤ gradient magnitude < 60, amount = 1.0; Low complexity region: Gradient magnitude < 30, amount = 0.

6. A method for adaptive real-time image quality enhancement, based on the system according to any one of claims 1-5, characterized in that, The method includes the following steps: Step S1: Convert the YUV format data of the client video stream to RGB format on the GPU side and further convert it to the HSL color space; Step S2: Calculate the non-linear enhancement coefficient α(L) according to the current frame luminance component L, enhance the luminance of the region where L ≤ 0.8, and limit the enhanced luminance value to L' ≤ 1; Step S3: Detect skin-color regions with chromaticity H ∈ [0°, 36°] and saturation S ∈ [0.2, 0.6] in the HSL space, and generate an optimized skin-color mask by combining the detection results of the adjacent 3 previous frames and 3 subsequent frames; Step S4: Generate a saturation gain coefficient matrix for non-skin-color regions through a pre-trained CNN model and perform local enhancement using a traditional saturation adjustment algorithm; Step S5: Divide the image into several blocks, calculate the variance σ² of each block, perform wavelet decomposition on the blocks with σ² ≥ 100, perform one-dimensional bilateral filtering on the low-frequency signal along the horizontal or vertical direction, and perform soft-threshold noise reduction on the high-frequency signal; Step S6: Calculate the image gradient magnitude, divide the high, medium, and low texture complexity regions, and dynamically adjust the sharpening intensity amount according to the gradient magnitude; Step S7: Inversely convert the processed HSL data to RGB format and perform real-time rendering output through the GPU shader pipeline, with a frame rate ≥ 30fps.

7. The method according to claim 6, characterized in that: In the step S3, the optimization of the skin-color mask is calculated by the following formula: Where, Mt: The skin-color region mask preliminarily detected in the t-th frame, with a value of 0 for non-skin-color or 1 for skin-color; Mt−3, Mt−2, …, Mt+3: The skin-color masks of the adjacent 3 previous frames and 3 subsequent frames; Denominator 7: The total number of frames participating in the joint optimization, the current frame + 3 frames before and after; Mt is the skin-color region preliminarily detected in the t-th frame, and when the optimized mask value ≥ 0.6, it is determined as a skin-color region.

8. The method according to claim 6, characterized in that: In the step S5, the formula for soft-threshold processing is: Where, x: The wavelet coefficient of the high-frequency signal; T: Dynamic threshold, positively correlated with the noise intensity, T = 15 when the noise intensity ≥ 60; sign(x): Take the sign of x, +1 or -1.

9. The method according to claim 6, characterized in that: In the step S6, the dynamic adjustment of the sharpening intensity satisfies: amount = 0.025 × gradient_magnitude And the maximum value of amount is limited to 1.5; gradient_magnitude: the magnitude of the image gradient, calculated by the Sobel operator.

10. The method according to claim 6, wherein: In the steps S1 to S7, the parallel computing on the GPU side is implemented through the following optimizations: The YUV→RGB conversion uses the fragment shader of OpenGL ES 3.0; The one-dimensional bilateral filtering is processed in parallel in blocks by the GPU compute shader Compute Shader; The rendering output adopts a double-buffer mechanism with a latency ≤ 30 ms.

Citation Information

Cited By

  • Video skin beautifying method and system fusing multi-color space and filtering

    CN121214522A

  • Video quality diagnosis method based on standardized color conversion and multistage adaptive detection

    CN122437916A