IPTV Video Quality Enhancement Methods and Systems
By performing inter-frame separation, adaptive color gamut mapping, and tensor decomposition optimization on IPTV video streams, the problem of ignoring inter-frame temporal correlation is solved, generating high-quality, low-complexity video enhancement streams and improving the motion consistency and visual realism of the video.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-26
- Publication Date
- 2026-04-03
AI Technical Summary
Existing video enhancement technologies ignore the temporal correlation between frames, resulting in flickering, jumping, or detail distortion during motion. They also have high computational complexity, making it difficult to meet the real-time and resource requirements of IPTV systems.
By performing inter-frame separation on IPTV video streams, extracting inter-frame temporal features, performing adaptive color gamut mapping and tensor decomposition, optimizing multi-dimensional texture feature maps, and combining high bit-depth video data streams for inter-frame dynamic compensation synthesis, a quality-enhanced video output stream is generated.
It achieves improved temporal consistency and visual comfort in video during motion, significantly enhancing the realism and visual effects of video footage while reducing computational complexity.
Smart Images

Figure CN121415329B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of video quality technology, and in particular to a method and system for enhancing IPTV video quality. Background Technology
[0002] Current mainstream video enhancement technologies mostly focus on optimizing the quality of single frames, lacking systematic processing of the overall temporal consistency and color continuity of the video. Although some studies have attempted to introduce techniques such as color gamut mapping and super-resolution reconstruction to improve image quality, they often ignore the temporal correlation between frames, leading to problems such as flickering, skipping, or detail distortion in the enhanced video during motion. In addition, traditional methods have high computational complexity when processing high bit-depth video data, making it difficult to meet the stringent real-time and resource requirements of IPTV systems. Summary of the Invention
[0003] The main objective of this invention is to provide an IPTV video quality enhancement method that solves the technical problem that existing technologies often ignore the temporal correlation between frames, resulting in flickering, skipping, or detail distortion in the enhanced video during motion.
[0004] To achieve the above objectives, the present invention provides an IPTV video quality enhancement method, comprising the following steps:
[0005] Inter-frame separation is performed on the IPTV video stream to obtain the video frame sequence and inter-frame temporal characteristics;
[0006] Based on the inter-frame temporal features, adaptive color gamut mapping is performed on the video frame sequence to obtain color reconstruction sequence data;
[0007] Tensor decomposition and recombination are performed on the color reconstruction sequence data to obtain a multi-dimensional texture feature map;
[0008] Dynamic bit-depth optimization is performed on the multi-dimensional texture feature map to obtain a high bit-depth video data stream;
[0009] Based on the high bit depth video data stream, inter-frame dynamic compensation synthesis is performed to obtain a quality-enhanced video output stream.
[0010] Furthermore, the step of performing inter-frame separation on the IPTV video stream to obtain the video frame sequence and inter-frame temporal features includes:
[0011] The IPTV video stream is subjected to temporal signal sampling analysis to obtain a video sampling data stream, and the video sampling data stream is subjected to frame synchronization mark extraction to obtain a frame synchronization feature matrix.
[0012] Frame boundary localization is performed on the frame synchronization feature matrix to obtain a video frame segmentation sequence, and inter-frame motion vector calculation is performed based on the video frame segmentation sequence to obtain inter-frame displacement feature vectors;
[0013] Based on the inter-frame displacement feature vector, the video frame segmentation sequence is temporally correlated and grouped to obtain a video frame sequence, and temporal features are extracted from the video frame sequence to obtain inter-frame temporal features.
[0014] Furthermore, the step of performing adaptive color gamut mapping on the video frame sequence based on the inter-frame temporal features to obtain color reconstruction sequence data includes:
[0015] Color component decoupling analysis is performed on the inter-frame temporal features to obtain the RGB three-channel component mapping matrix, and color space coordinate transformation is performed based on the RGB three-channel component mapping matrix to obtain the YUV color space feature parameters.
[0016] The YUV color gamut feature parameters are nonlinearly stretched by chromaticity histogram equalization to obtain the color distribution equalization coefficient, and chromaticity saturation is corrected based on the color distribution equalization coefficient to obtain the chromaticity correction mapping table.
[0017] The video frame sequence is remapped for luminance channel based on the chroma correction mapping table to obtain a luminance enhancement sequence, and then chroma channel compensation is performed on the luminance enhancement sequence to obtain a chroma compensation vector.
[0018] The chromaticity compensation vector is dynamically mapped by chromaticity gamma calibration to obtain color reconstruction sequence data, which includes chromaticity component reconstruction coefficients, luminance mapping curves, and chromaticity calibration parameter sets.
[0019] Furthermore, the tensor decomposition and recombination of the color reconstruction sequence data to obtain a multi-dimensional texture feature map includes:
[0020] The color reconstruction sequence data is constructed using a third-order tensor to obtain a spatiotemporal color feature tensor, and the nuclear norm minimization decomposition is performed based on the spatiotemporal color feature tensor to obtain the tensor core coefficient matrix.
[0021] The tensor core coefficient matrix is subjected to multimodal separation by Tucker decomposition to obtain three subspace projection matrices in the time domain, spatial domain, and color domain. Feature correlation analysis is then performed on the three subspace projection matrices to obtain a cross-modal feature correlation map.
[0022] Texture features are extracted based on the cross-modal feature association map to obtain multi-scale texture descriptors, and the multi-scale texture descriptors are spatially reorganized to obtain a texture structure reconstruction matrix.
[0023] The texture structure reconstruction matrix is dimensionally fused by singular value constraints to obtain a multi-dimensional texture feature map, which includes a texture edge gradient map, a structural saliency map, and a texture orientation field intensity distribution map.
[0024] Furthermore, the step of dynamically optimizing the multi-dimensional texture feature map to obtain a high-bit-depth video data stream includes:
[0025] Bit depth quantization analysis is performed on the multi-dimensional texture feature map to obtain a quantization error distribution map, and local bit depth statistics are performed based on the quantization error distribution map to obtain a bit depth distribution histogram.
[0026] The bit depth distribution histogram is segmented and quantized using a non-uniform quantizer to obtain a multi-level quantization parameter set. Local bit allocation is then performed on the multi-level quantization parameter set to obtain a bit allocation weight table.
[0027] Based on the bit allocation weight table, adaptive quantization compensation is performed on the texture region in the multi-dimensional texture feature map to obtain the texture compensation coefficient matrix, and bit depth optimization reconstruction is performed on the texture compensation coefficient matrix to obtain the bit depth reconstruction mapping table.
[0028] The video stream is reassembled by dynamically mapping the bit depth to obtain a high bit depth video data stream, which includes quantization compensation coefficients, bit depth optimization parameters, and bit depth mapping curves.
[0029] Furthermore, the step of performing inter-frame dynamic compensation synthesis based on the high bit-depth video data stream to obtain a quality-enhanced video output stream includes:
[0030] Dense optical flow field estimation is performed based on the high bit depth video data stream to obtain a motion vector distribution map, and motion smoothing filtering is applied to the motion vector distribution map to obtain a temporally consistent motion vector field;
[0031] Inter-frame residual prediction is performed on the high bit-depth video data stream using the time-consistent motion vector field to obtain a predicted residual map, and adaptive compensation coefficients are generated based on the predicted residual map to obtain a compensation coefficient matrix.
[0032] Based on the compensation coefficient matrix, the high bit depth video data stream is spatiotemporally fused and synthesized to obtain a synthesized video frame sequence. Based on the synthesized video frame sequence, multi-scale detail enhancement is performed to obtain a detail-enhanced video stream.
[0033] The detail-enhanced video stream is compressed using entropy coding optimization to obtain a compressed video data stream. Then, quality-adaptive filtering is performed on the compressed video data stream to obtain a quality-enhanced video output stream.
[0034] Furthermore, the step of performing spatiotemporal fusion synthesis on the high bit-depth video data stream based on the compensation coefficient matrix to obtain a synthesized video frame sequence includes:
[0035] Multi-scale spatial feature extraction is performed on the high bit-depth video data stream to obtain a spatial gradient feature map, and image segmentation is performed based on the spatial gradient feature map to obtain a video content region mask.
[0036] The video content region mask is adaptively weighted using the compensation coefficient matrix to obtain a dynamic fusion weight map, and the high bit depth video data stream is pixel-level compensated based on the dynamic fusion weight map to obtain a compensated pixel mapping table.
[0037] Based on the compensated pixel mapping table, the high bit-depth video data stream is subjected to spatiotemporal domain fusion processing to obtain a spatiotemporally aligned video frame sequence, and edge-aware filtering is performed on the spatiotemporally aligned video frame sequence to obtain a smoothly transitioned video frame sequence.
[0038] The smooth transition video frame sequence is reconstructed using Laplacian pyramids to perform multi-resolution synthesis, resulting in a high-resolution video frame sequence. Color consistency correction is then performed on the high-resolution video frame sequence to obtain a synthesized video frame sequence.
[0039] The present invention also provides an IPTV video quality enhancement system, comprising:
[0040] The separation module is used to perform inter-frame separation of IPTV video streams to obtain video frame sequences and inter-frame temporal characteristics;
[0041] The mapping module is used to perform adaptive color gamut mapping on the video frame sequence based on the inter-frame temporal features to obtain color reconstruction sequence data;
[0042] The recombination module is used to perform tensor decomposition and recombination on the color reconstruction sequence data to obtain a multi-dimensional texture feature map.
[0043] The optimization module is used to perform dynamic bit-depth optimization on the multi-dimensional texture feature map to obtain a high bit-depth video data stream;
[0044] The synthesis module is used to perform inter-frame dynamic compensation synthesis based on the high bit depth video data stream to obtain a quality-enhanced video output stream.
[0045] The present invention also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of any of the methods described above.
[0046] The present invention also provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the steps of any of the methods described above.
[0047] This invention provides an IPTV video quality enhancement method, comprising the following steps: performing inter-frame separation on the IPTV video stream to obtain a video frame sequence and inter-frame temporal features; performing adaptive color gamut mapping on the video frame sequence based on the inter-frame temporal features to obtain color reconstruction sequence data; performing tensor decomposition and recombination on the color reconstruction sequence data to obtain a multi-dimensional texture feature map; performing dynamic bit-depth optimization on the multi-dimensional texture feature map to obtain a high bit-depth video data stream; and performing inter-frame dynamic compensation synthesis based on the high bit-depth video data stream to obtain a quality-enhanced video output stream. This method solves the technical problem that existing technologies often ignore the temporal correlation between frames, leading to flickering, skipping, or detail distortion in the enhanced video during motion. It achieves adaptive color gamut mapping based on inter-frame temporal features, which can dynamically adjust the color distribution according to different scene content, avoiding the color distortion problem caused by traditional fixed mapping methods, and significantly improving the realism and visual comfort of the video picture. Attached Figure Description
[0048] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0049] Figure 1 This is a schematic diagram of the steps of an IPTV video quality enhancement method in one embodiment of the present invention;
[0050] Figure 2 This is a structural block diagram of an IPTV video quality enhancement system according to an embodiment of the present invention;
[0051] Figure 3 This is a schematic block diagram of the structure of a computer device according to an embodiment of the present invention.
[0052] The objectives, features, and advantages of this invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0053] The embodiments of the present invention are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and should not be construed as limiting the present invention. The step numbers in the following embodiments are set only for ease of explanation, and there is no limitation on the order between the steps. The execution order of each step in the embodiments can be adaptively adjusted according to the understanding of those skilled in the art.
[0054] The following describes in detail, with reference to the accompanying drawings, an IPTV video quality enhancement method according to an embodiment of the present invention. First, the IPTV video quality enhancement method according to an embodiment of the present invention will be described with reference to the accompanying drawings.
[0055] Figure 1 This invention provides an IPTV video quality enhancement method according to one embodiment, comprising the following steps:
[0056] Step S1: Perform inter-frame separation on the IPTV video stream to obtain the video frame sequence and inter-frame temporal characteristics.
[0057] Specifically, in the process of inter-frame separation of IPTV video streams, the main step is to analyze the video bitstream during transmission, decompose it into independent video frame sequences, and then extract inter-frame temporal features between adjacent frames to reflect the temporal correlation of object movement, scene switching, and image transitions in the video content. This step, as the foundation of the entire video quality enhancement process, can typically be implemented using motion estimation-based analysis methods or optical flow algorithms. For example, in the MPEG series coding standards, the inter-frame prediction mechanism naturally includes temporal relationship information between frames. By analyzing the motion vectors carried by P-frames and B-frames, inter-frame temporal features for subsequent processing can be obtained. For instance, in an IPTV live sports event application scenario, due to the rapid and frequent movement of athletes and drastic changes in the image, without effective inter-frame separation and temporal feature extraction, problems such as color jumps and motion blur may occur during subsequent color reconstruction. This step accurately identifies the motion trends between frames, providing a reliable temporal basis for adaptive color gamut mapping, ensuring that the enhanced video has a more natural and smooth dynamic performance, and improving the consistency and comfort of the overall viewing experience.
[0058] Step S2: Based on the inter-frame temporal features, perform adaptive color gamut mapping on the video frame sequence to obtain color reconstruction sequence data.
[0059] Specifically, adaptive color gamut mapping of the video frame sequence based on the inter-frame temporal features refers to dynamically adjusting the color distribution of each frame image based on the extracted inter-frame motion trends and temporal continuity information to achieve color reconstruction sequence data that better matches the original scene or display device characteristics. This step inputs the inter-frame temporal features as control parameters into the color gamut mapping algorithm, making the color conversion process no longer a static processing of isolated frames, but capable of real-time adjustment based on the changing trends between adjacent frames. This avoids color jumps and distortions caused by rapid scene switching or intense motion. For example, in IPTV live sports broadcasts, when cameras quickly switch shots or players run at high speed, there are significant motion differences between video frames. If a traditional fixed color gamut mapping method is used, it can easily cause color abrupt changes or artifacts. However, by introducing an adaptive mechanism driven by inter-frame temporal features, the system can identify the color transition relationship between the current frame and the preceding and following frames, and optimize the color mapping path accordingly, ensuring a smooth color transition in the temporal dimension. This improves the overall visual coherence and realism of the video, providing a high-quality color foundation for subsequent texture enhancement and high-bit-depth compositing.
[0060] Step S3: Perform tensor decomposition and recombination on the color reconstruction sequence data to obtain a multi-dimensional texture feature map.
[0061] Specifically, tensor decomposition and recombination of the color reconstruction sequence data refers to modeling the video frame sequence after adaptive color gamut mapping as a multidimensional array (i.e., tensor), and then decomposing it into multiple low-rank subspace components using mathematical decomposition methods (such as CP decomposition or Tucker decomposition), thereby extracting the potential structural features of the video in multiple dimensions such as space, color, and time. Subsequently, by purposefully recombinating these low-dimensional features, a multidimensional texture feature map that reflects the local texture details and global structural information of the video is generated. This process not only preserves the spatial texture information of the video content but also integrates inter-frame temporal changes and color distribution features, enabling subsequent processing to more accurately identify and enhance image details under different motion states. For example, in IPTV live sports broadcasts, the fine stripes on players' jerseys, the complex textures of the stands, and the blurred edges in fast-moving scenes all place high demands on texture representation capabilities. Tensor decomposition and reconstruction techniques can effectively separate these different levels of texture structures and enhance and reconstruct them in multi-dimensional feature maps, thereby providing richer and more accurate visual information support for dynamic bit depth optimization and improving the clarity and realism of videos in complex motion scenes.
[0062] Step S4: Perform dynamic bit-depth optimization on the multi-dimensional texture feature map to obtain a high bit-depth video data stream.
[0063] Specifically, dynamic bit-depth optimization of the multi-dimensional texture feature map refers to adaptively adjusting the quantization precision of pixel values based on the visual sensitivity and content complexity of different regions, after extracting multi-dimensional texture features such as video space, color, and time. This transforms the original video data into a video data stream with higher bit depth. This process dynamically allocates bit-depth enhancement strategies by analyzing gradient changes, edge information, and motion intensity in local regions of the texture feature map. For example, higher bit depth is used in areas with rich detail or intense motion, while lower bit depth is maintained in smooth areas to save resources, ultimately achieving a balance between image quality improvement and bandwidth efficiency. For instance, in IPTV live sports broadcasts, blurred edges caused by rapid athlete movement, dense textures in the stands, and changes in the gloss of the stadium surface all pose challenges to the dynamic range of the image. Through this step, the system can enhance the smoothness of grayscale transitions and color gradation while preserving this key visual information. This results in a more realistic lighting effect and stronger visual immersion on HDR display devices, significantly improving the user's viewing experience.
[0064] Step S5: Perform inter-frame dynamic compensation synthesis based on the high bit depth video data stream to obtain a quality-enhanced video output stream.
[0065] Specifically, inter-frame dynamic compensation synthesis based on the high bit-depth video data stream refers to, after completing the high bit-depth conversion, combining the motion information and temporal continuity characteristics between adjacent frames in the video sequence to finely adjust the brightness, color, and texture transitions between frames, thereby generating a visually coherent and dynamically natural enhanced video output stream. This step utilizes the inter-frame temporal features extracted in the preprocessing stage and the rich grayscale information contained in the high bit-depth data, employing motion-compensated interpolation algorithms or optical flow estimation techniques to dynamically optimize the transition regions between frames, reducing edge blurring and detail distortion caused by motion estimation errors or bit-depth increases. For example, in IPTV live sports broadcasts, when players are running at high speed or the camera is switching rapidly, the content changes drastically between frames. Without dynamic compensation and synthesis, this can easily lead to image jumps, ghosting, or flickering. However, through this step, the system can reconstruct more natural and smooth intermediate frames or correct inter-frame transitions based on the fine textures and motion trends in the high bit depth video data stream. This achieves high-quality consistency of the enhanced video output stream in the time dimension, significantly improving the realism and immersion of the user's viewing experience.
[0066] In a specific embodiment, the step of performing inter-frame separation on the IPTV video stream to obtain the video frame sequence and inter-frame temporal characteristics includes:
[0067] The IPTV video stream is subjected to temporal signal sampling analysis to obtain a video sampling data stream, and the video sampling data stream is subjected to frame synchronization mark extraction to obtain a frame synchronization feature matrix.
[0068] Frame boundary localization is performed on the frame synchronization feature matrix to obtain a video frame segmentation sequence, and inter-frame motion vector calculation is performed based on the video frame segmentation sequence to obtain inter-frame displacement feature vectors;
[0069] Based on the inter-frame displacement feature vector, the video frame segmentation sequence is temporally correlated and grouped to obtain a video frame sequence, and temporal features are extracted from the video frame sequence to obtain inter-frame temporal features.
[0070] Specifically, the IPTV video stream is first subjected to temporal signal sampling analysis. This involves collecting video data packets segment by segment according to the timestamp information defined in standard video coding protocols (such as H.264 / AVC or H.265 / HEVC), and recording the arrival time of each frame in conjunction with a time baseline, thereby generating a video sampling data stream. For example, in a 1080p@30fps high-definition live stream, the system collects 30 complete video frames per second and their corresponding transmission timestamps, forming a data stream with precise time resolution. Subsequently, the system extracts frame synchronization markers from the video sampling data stream. This involves identifying key syntax elements such as the NAL unit start code, SPS / PPS parameter set, and frame header identifier to extract synchronization information representing frame type (I / P / B frames) and frame sequence number. This results in the construction of a frame synchronization feature matrix, where each row corresponds to the synchronization metadata of a frame, including fields such as frame type, display order, and decoding order. Building upon this, the system further performs frame boundary localization on the frame synchronization feature matrix. Specifically, it uses the frame start position information in the frame synchronization markers to slice and reassemble the video data stream, accurately separating each video frame from the continuous data stream to form a video frame segmentation sequence. For example, in a GOP structure containing multiple B-frames, the system can accurately locate the start byte offset between each P-frame and B-frame based on the synchronization markers, ensuring that subsequent frame processing will not result in misalignment or misreading. Subsequently, the system performs inter-frame motion vector calculation based on the video frame segmentation sequence. This involves using block matching algorithms or optical flow estimation methods between adjacent frames to calculate the spatial displacement direction and magnitude of pixel blocks, thereby generating inter-frame displacement feature vectors. These vectors not only contain horizontal and vertical displacement values (e.g., dx=12, dy=-7) but also record the displacement intensity distribution and motion pattern complexity, providing a basis for subsequent temporal correlation grouping. Next, the system performs temporal correlation grouping on the video frame segmentation sequence based on the inter-frame displacement feature vector. That is, based on the motion similarity and temporal continuity between frames, highly correlated frames are grouped together to generate a video frame sequence. For example, in a live football match, the system identifies numerous high-frequency motion regions caused by players running quickly and classifies several consecutive frames into a motion-active group; while shots with relatively static backgrounds are classified as static. This grouping strategy helps subsequent processing modules select appropriate enhancement strategies based on different groups. Finally, the system extracts temporal features from the video frame sequence, extracting the temporal variation patterns from each group of frames, including frame rate fluctuations, brightness change trends, and texture dynamic range expansion, forming inter-frame temporal features.For example, in a program segment with frequent shot changes, the system detects that the frame rate jumps from 25fps to 50fps, while the average brightness value fluctuates by more than 15% between adjacent frames. This information is encapsulated in a set of temporal features as an important reference for subsequent adaptive enhancement and bit rate control.
[0071] In a specific embodiment, the step of performing adaptive color gamut mapping on the video frame sequence based on the inter-frame temporal features to obtain color reconstruction sequence data includes:
[0072] Color component decoupling analysis is performed on the inter-frame temporal features to obtain the RGB three-channel component mapping matrix, and color space coordinate transformation is performed based on the RGB three-channel component mapping matrix to obtain the YUV color space feature parameters.
[0073] The YUV color gamut feature parameters are nonlinearly stretched by chromaticity histogram equalization to obtain the color distribution equalization coefficient, and chromaticity saturation is corrected based on the color distribution equalization coefficient to obtain the chromaticity correction mapping table.
[0074] The video frame sequence is remapped for luminance channel based on the chroma correction mapping table to obtain a luminance enhancement sequence, and then chroma channel compensation is performed on the luminance enhancement sequence to obtain a chroma compensation vector.
[0075] The chromaticity compensation vector is dynamically mapped by chromaticity gamma calibration to obtain color reconstruction sequence data, which includes chromaticity component reconstruction coefficients, luminance mapping curves, and chromaticity calibration parameter sets.
[0076] Specifically, the process of adaptively mapping the video frame sequence based on the inter-frame temporal features to obtain color reconstruction sequence data is a key technical step in the entire IPTV video quality enhancement method to achieve color consistency and improve visual comfort. The core of this step lies in combining motion information in the temporal dimension with color space processing. Through multi-stage processing such as dynamic analysis of the RGB three channels, YUV color gamut conversion, histogram equalization, collaborative optimization of luminance and chrominance channels, and gamma calibration, the final output is color reconstruction sequence data containing chrominance component reconstruction coefficients, luminance mapping curves, and a color gamut calibration parameter set, thus providing a high-quality color foundation for subsequent texture enhancement and high-bit-depth synthesis. Specifically, the system first performs color component decoupling analysis on the inter-frame temporal features, that is, by analyzing the color change trends between adjacent frames, it identifies color channels (such as R, G, B) that remain stable or undergo significant changes between different frames. For example, in live sports broadcasts, the color of a player's jersey might slightly shift or flicker during rapid movement. In this case, the system uses inter-frame temporal characteristics to determine which color channel changes are due to actual content changes and which are caused by encoding compression or transmission noise, and constructs an RGB three-channel component mapping matrix accordingly. This matrix reflects the color relationship between each channel in each frame and the preceding and following frames, providing a basis for subsequent color space coordinate transformation. Next, the system performs color space coordinate transformation based on the RGB three-channel component mapping matrix, mapping pixel values in the RGB color space to the YUV color space, thereby obtaining YUV color space feature parameters. The advantage of the YUV format is that it separates luminance (Y) and chrominance (U, V), allowing subsequent processing to optimize luminance and chrominance separately, avoiding color interference problems. Subsequently, the system uses chrominance histogram equalization to non-linearly stretch the YUV color space feature parameters to improve the overall color distribution uniformity of the image. For example, in some low-light or high-contrast shots, the U and V channels may be concentrated, leading to insufficient color saturation or local distortion. At this point, the system stretches the chroma histograms of the U and V channels, expanding the chroma values that were originally concentrated in a small range to a wider range, thereby improving the color gradation and expressiveness. Based on this, the system further calculates a color distribution balance coefficient and performs chroma saturation correction based on this coefficient, generating a chroma correction mapping table. This mapping table records the correction target value corresponding to each original chroma value, ensuring that oversaturation or artifacts do not occur after color enhancement. For example, in nighttime broadcasts of football matches, green grass may appear dull or yellowish due to uneven lighting. After chroma correction, the system can automatically identify and enhance the saturation of green hues, making it appear a more natural grass green. Next, the system remaps the luminance channels of the video frame sequence based on the chroma correction mapping table to obtain a luminance-enhanced sequence.This process primarily utilizes information from the Y channel (i.e., the luminance channel) for non-linear adjustments to enhance image contrast and detail. For example, during rapid shot transitions, exposure changes may cause some areas to be too bright or too dark. In this case, the system dynamically adjusts the luminance mapping curve based on the current frame's Y channel statistics (such as average luminance value, maximum and minimum values), combined with the continuity requirements of luminance transitions in inter-frame temporal characteristics. This ensures that the enhanced luminance distribution meets the content requirements of the current frame while maintaining a smooth transition with preceding and following frames. Simultaneously, to avoid chroma drift caused by luminance adjustments, the system also performs chroma channel compensation on the luminance enhancement sequence, obtaining a chroma compensation vector. For example, when the luminance of a frame is increased, skin tones that were originally in the midtones may become slightly reddish or bluish. The system calculates the appropriate chroma compensation direction and magnitude based on the chroma information of historical frames to ensure that the enhanced image colors remain natural. Finally, the system dynamically maps the chroma compensation vector through gamut gamma calibration to adapt to the color gamut characteristics of different display devices, thereby ultimately generating color reconstruction sequence data. The gamma calibration process remaps color data based on the display characteristics of different terminals (e.g., HDR displays typically use PQ gamma curves, while traditional SDR displays use sRGB gamma curves), ensuring that the enhanced video presents optimal results regardless of the device on which it is played. For example, when watching sports events on an HDR-enabled smart TV, the system expands luminance and chrominance according to the HDR10 standard, making the green of the grass on the field more vibrant and the blue of the sky more transparent; while on a regular mobile phone screen, the dynamic range is appropriately compressed to avoid overexposure. The final output color reconstruction sequence data includes chrominance component reconstruction coefficients (used to describe the enhancement ratio of each chrominance channel), luminance mapping curves (used to control the intensity of luminance enhancement), and a color gamut calibration parameter set (used to adapt to the color spaces of different display devices). In summary, this step achieves refined color reconstruction of video frame sequences through several key technical steps, including color component decoupling analysis driven by inter-frame temporal characteristics, color gamut space conversion, chrominance histogram equalization, luminance and chrominance channel co-optimization, and gamma dynamic mapping. In the entire IPTV video stream processing workflow, this step not only effectively solves problems such as color jumps and saturation imbalances under the traditional fixed color gamut mapping method, but also ensures the color consistency and visual comfort of the enhanced video in the time dimension by introducing an inter-frame dynamic compensation mechanism. This provides high-quality input data support for subsequent tensor texture enhancement and high bit depth optimization, thereby comprehensively improving the overall picture quality of the video.
[0077] In a specific embodiment, the step of performing tensor decomposition and recombination on the color reconstruction sequence data to obtain a multi-dimensional texture feature map includes:
[0078] The color reconstruction sequence data is constructed using a third-order tensor to obtain a spatiotemporal color feature tensor, and the nuclear norm minimization decomposition is performed based on the spatiotemporal color feature tensor to obtain the tensor core coefficient matrix.
[0079] The tensor core coefficient matrix is subjected to multimodal separation by Tucker decomposition to obtain three subspace projection matrices in the time domain, spatial domain, and color domain. Feature correlation analysis is then performed on the three subspace projection matrices to obtain a cross-modal feature correlation map.
[0080] Texture features are extracted based on the cross-modal feature association map to obtain multi-scale texture descriptors, and the multi-scale texture descriptors are spatially reorganized to obtain a texture structure reconstruction matrix.
[0081] The texture structure reconstruction matrix is dimensionally fused by singular value constraints to obtain a multi-dimensional texture feature map, which includes a texture edge gradient map, a structural saliency map, and a texture orientation field intensity distribution map.
[0082] Specifically, the process of tensor decomposition and recombination of the color reconstruction sequence data to obtain a multi-dimensional texture feature map is a crucial technical step in the entire IPTV video quality enhancement method, aimed at improving the ability to express video content details and perceive structure. This step involves constructing a high-order tensor model, performing mathematical tensor decomposition and multimodal separation, extracting and organizing multi-scale texture information, and finally fusing it to generate a multi-dimensional texture feature map containing key visual information such as edge gradients, structural saliency, and directional field intensity. This provides rich and accurate image semantic support for subsequent dynamic bit-depth optimization. In the specific implementation process, the system first constructs a third-order tensor for the color reconstruction sequence data. This involves arranging each frame of the image after adaptive color gamut mapping in chronological order and preserving its spatial two-dimensional structure (such as width × height) and color channel information (such as RGB or YUV) within each frame, thereby forming a spatiotemporal color feature tensor with three dimensions: time, space, and color. For example, in a live sports video played at 1080p resolution and 30 frames per second, the tensor can be represented as a data structure of size T×H×W×C, where T represents the number of frames (e.g., 50 consecutively captured frames), H and W represent the height and width of the image (both 1920 pixels), and C represents the number of color channels (e.g., 3-channel RGB). This process not only abstracts the video content into a unified mathematical structure but also lays the foundation for subsequent tensor analysis. Subsequently, the system performs kernel norm minimization decomposition based on the spatiotemporal color feature tensor, that is, using the idea of low-rank approximation, extracting the tensor core coefficient matrix that reflects the overall structural characteristics from the original tensor. This operation is essentially a denoising and feature compression process, aiming to remove noise interference caused by transmission errors or encoding compression in the video while preserving the main texture and motion structure of the video content. For example, in a football match, blocky artifacts may appear when players run fast due to encoding distortion. In this case, kernel norm minimization decomposition will automatically identify and suppress these unstructured components, making the output core coefficient matrix more focused on describing the real motion trajectory and texture structure. Furthermore, the system performs multimodal separation on the tensor core coefficient matrix through Tucker decomposition, decomposing it into three independent but interrelated subspace projection matrices: spatial, temporal, and color. These three projection matrices correspond to spatial features (such as spatial structures like edges and corners), temporal features (such as motion vectors and inter-frame variation trends), and color gamut features (such as color distribution and saturation changes) in the video content, respectively. For example, in a scene where an athlete is running at high speed, the spatial projection matrix highlights the stripes on the jersey and the body outline, the temporal projection matrix captures the continuity of movement and speed changes, and the color gamut projection matrix reflects the trends in skin color and clothing color changes with lighting conditions.Based on this, the system performs feature correlation analysis on the projection matrices of these three subspaces to construct a cross-modal feature association map, which reveals the collaborative relationships between different modalities, such as how spatial edges evolve over time and how color changes affect motion perception. Next, the system extracts texture features based on the cross-modal feature association map to obtain multi-scale texture descriptors. This process mainly relies on convolutional filter banks or multi-scale transform algorithms (such as wavelet transform) to extract local texture information from different frequency levels in the image. For example, in a scene with a dense crowd in the audience, low-frequency parts may appear as large areas of dark gray, while high-frequency parts contain rich facial contours and expression details. By analyzing these multi-scale texture descriptors, the system can accurately identify which areas belong to the static background and which belong to the dynamic foreground, thus providing a basis for subsequent spatial organization. Subsequently, the system reorganizes the multi-scale texture descriptors in the spatial domain, that is, it re-integrates the texture information at each scale into a unified texture structure reconstruction matrix according to the spatial layout of the video frames. This process is similar to a "jigsaw puzzle," ensuring that each local texture can find its correct location in the global image, avoiding structural misalignment problems caused by viewpoint switching or motion blur. Finally, the system performs dimensionality fusion on the texture structure reconstruction matrix using singular value constraints to obtain a multi-dimensional texture feature map. Singular value constraints refer to compressing or adjusting secondary components while keeping the principal components of the matrix unchanged, thereby improving computational efficiency and enhancing the stability of the image structure. For example, in high-speed motion scenes, some edges may experience slight jitter due to inter-frame inconsistencies. In this case, the singular value constraint mechanism automatically smooths out these unstable factors, making the edges clearer and more stable. The final generated multi-dimensional texture feature map includes a texture edge gradient map (used to describe the changes in edge intensity in the image; higher values indicate clearer edges), a structural saliency map (used to identify areas of visual attention in the image, such as player faces, goal frames, etc.), and a texture orientation field intensity distribution map (used to characterize the directional consistency of the texture, such as the stripe direction of grass, the arrangement direction of spectator seats, etc.). In summary, this step utilizes several key technologies, including third-order tensor modeling, nuclear norm minimization decomposition, Tucker multimodal separation, cross-modal feature correlation analysis, multi-scale texture extraction and organization, and singular value constraint-driven dimensional fusion, to achieve deep structural mining and texture enhancement of the video content after color reconstruction. In the entire IPTV video quality enhancement process, this step not only effectively improves the texture restoration capability of the video in complex motion scenes but also provides more refined image semantic support for subsequent high-depth optimization and inter-frame dynamic compensation by introducing multi-dimensional feature maps. Especially in typical application scenarios with frequent dynamic scenes and high detail requirements, such as live sports broadcasts, this tensor decomposition and reconstruction mechanism exhibits stronger structure awareness and texture preservation capabilities, thereby significantly improving the overall expressiveness and realism of the video quality.
[0083] In a specific embodiment, the step of dynamically optimizing the bit depth of the multi-dimensional texture feature map to obtain a high bit depth video data stream includes:
[0084] Bit depth quantization analysis is performed on the multi-dimensional texture feature map to obtain a quantization error distribution map, and local bit depth statistics are performed based on the quantization error distribution map to obtain a bit depth distribution histogram.
[0085] The bit depth distribution histogram is segmented and quantized using a non-uniform quantizer to obtain a multi-level quantization parameter set. Local bit allocation is then performed on the multi-level quantization parameter set to obtain a bit allocation weight table.
[0086] Based on the bit allocation weight table, adaptive quantization compensation is performed on the texture region in the multi-dimensional texture feature map to obtain the texture compensation coefficient matrix, and bit depth optimization reconstruction is performed on the texture compensation coefficient matrix to obtain the bit depth reconstruction mapping table.
[0087] The video stream is reassembled by dynamically mapping the bit depth to obtain a high bit depth video data stream, which includes quantization compensation coefficients, bit depth optimization parameters, and bit depth mapping curves.
[0088] Specifically, the process of dynamically optimizing the bit depth of the multi-dimensional texture feature map to obtain a high bit depth video data stream is a key technical step in the entire IPTV video quality enhancement method to achieve dynamic range expansion and visual detail enhancement of the image. This step involves quantization analysis, segmented encoding, bit allocation, and adaptive compensation of the local texture characteristics of the video content, ultimately generating a high bit depth video data stream containing quantization compensation coefficients, bit depth optimization parameters, and bit depth mapping curves, thereby maximizing image quality improvement under limited bandwidth conditions. In the specific implementation process, the system first performs bit depth quantization analysis on the multi-dimensional texture feature map, that is, based on the pixel value change trends in the texture edge gradient map, structural saliency map, and texture direction field intensity distribution map, it assesses the bit depth requirements of the current image in different regions and constructs a quantization error distribution map accordingly. For example, in the application scenario of live sports broadcasts, the fine stripe areas on the player's jersey contain high-frequency texture information. If a fixed 8-bit quantization method is used, color banding or false contour phenomena are likely to occur. However, through quantization error distribution map analysis, it can be found that the error values in these areas are significantly higher than those in the background areas, indicating that they require higher bit depth representation accuracy. Subsequently, the system performs local bit depth statistics based on the quantization error distribution map, calculates the minimum effective bit depth required in each local region, and generates a bit depth distribution histogram to describe the bit depth requirement distribution of each region in the entire image. For example, statistics show that 60% of the image region can meet visual requirements with an 8-bit bit depth, while 25% of the region needs to be increased to 10 bits, and the remaining 15% of high-texture regions even require bit depth support of 12 bits or more. Further, the system performs segmented quantization encoding on the bit depth distribution histogram using a non-uniform quantizer, that is, dividing the entire image into multiple sub-regions with different quantization step sizes according to the bit depth requirements of each region, and assigning a corresponding multi-level quantization parameter set to each sub-region. Compared with the traditional uniform quantization method, this non-uniform quantization strategy can reasonably control the growth of the overall data volume while ensuring the quality of key visual regions. For example, in a high-definition frame with a resolution of 1920×1080, the system might divide the entire image into 32 sub-blocks (e.g., each block is 480×270 pixels), and set a quantization step size for each sub-block: for the area containing the athlete's face, a quantization step size of 1.2 gray levels per unit is set, while for low-texture areas such as the audience seating area, a quantization step size of 3.0 gray levels per unit is set. This preserves key details while saving encoding resources. Subsequently, the system performs local bit allocation on the multi-level quantization parameter set, that is, dynamically adjusts the bit budget occupied by each sub-block according to the importance of different regions and their quantization error levels, generating a bit allocation weight table.For example, for regions with rich texture and large errors, the system allocates more bit resources, such as expanding an original 8-bit region to 10 bits, equivalent to adding 2 bits per pixel, increasing the total number of bits in that region by approximately 25%. For smooth regions, bit usage is appropriately compressed to ensure the overall bit rate remains within the allowable transmission bandwidth. Next, the system performs adaptive quantization compensation on the texture regions in the multi-dimensional texture feature map based on the bit allocation weight table. That is, while maintaining the original image structure, quantization error correction is applied to each sub-block to obtain a texture compensation coefficient matrix. This matrix records the compensation value of each pixel relative to the original quantization result, used for subsequent bit-depth optimization and reconstruction. For example, in a shot of a soccer ball passing at high speed, the surface of the ball experiences discontinuous color transitions due to changes in lighting. In this case, the system applies a higher compensation coefficient (e.g., +12 to +18) to this region, making the brightness and chromaticity changes more subtle and avoiding "jump" phenomena. Subsequently, the system performs bit-depth optimization and reconstruction on the texture compensation coefficient matrix. This involves re-establishing a higher-precision mapping relationship using the compensated pixel values, forming a bit-depth reconstruction mapping table. This table contains the optimized output value corresponding to each original pixel value, ensuring accurate restoration of high-quality images in subsequent video compositing stages. Finally, the system uses dynamic bit-depth mapping to reassemble the bit-depth reconstruction mapping table into the video stream. This involves re-encapsulating the optimized pixel values into the video bitstream according to a new bit-depth format, ultimately generating a high-bit-depth video data stream. This data stream includes not only quantization compensation coefficients (describing the error correction magnitude of each pixel, such as an average compensation value of +15) and bit-depth optimization parameters (indicating the final bit-depth configuration of each region, such as increasing the bit depth from 8 bits to 10 bits in some regions), but also bit-depth mapping curves (defining the mapping relationship between different gray levels, such as using finer mapping intervals in dark areas to enhance contrast). For example, when playing this optimized video on an HDR display device, the system adapts the bit depth mapping curve according to the HDR10 standard, making the green of the grass on the field more vibrant, the blue of the sky more transparent, and the details of the players' movements clearly visible even in shadow areas. In summary, this step achieves refined bit depth optimization of multi-dimensional texture feature maps through several key technologies, including bit depth quantization analysis, non-uniform quantization coding, local bit allocation, adaptive quantization compensation, and bit depth reconstruction. In the entire IPTV video quality enhancement process, this step not only effectively solves problems such as detail loss and severe artifacts under traditional fixed bit depth processing methods, but also improves the dynamic performance and visual realism of the video in complex motion scenes by introducing a dynamic bit allocation mechanism.Especially in typical application scenarios like live sports broadcasts, where the visuals are dynamic and rich in texture details, this dynamic bit-depth optimization mechanism demonstrates stronger detail expression capabilities and resource utilization efficiency, thus significantly improving the overall performance and viewing experience of the video. It's important to note that high bit-depth video data streams refer to a data format that uses a higher bit depth than traditional methods to represent the brightness and color information of pixels in each frame of an image during video processing and transmission. Here, "high" refers to the higher number of bits used in each color channel compared to the conventional standard, resulting in more delicate and richer transitions between light and dark areas, color gradations, and details. Compared to traditional 8-bit video, high bit-depth video typically uses 10-bit or even 12-bit quantization, meaning that the number of gray levels that each color channel can represent increases from 256 levels to 1024 or 4096 levels, thereby greatly enhancing the dynamic range and color transition capabilities of the image. This high bit-depth data format not only more accurately reproduces the light and shadow changes in the original scene, but also effectively avoids visual artifacts such as color banding and jagged edges common in low bit-depth formats. Especially when the display device supports HDR (High Dynamic Range), the high bit-depth video data stream can present a picture effect that is closer to the real perception of the human eye. In practical applications, such as when IPTV broadcasts a football match, after the images captured by the camera have undergone multiple enhancement steps, the system will perform bit-depth expansion on complex content such as player jersey stripes, grass details, and the expressions of the crowd in the stands, based on local texture features and visually important areas. This upscals some areas of the original 8-bit YUV420 format to 10 bits or even higher, and the final generated video data stream is called a high bit-depth video data stream. In subsequent inter-frame dynamic compensation synthesis, such a data stream can retain more details and achieve more natural inter-frame transitions, so that the picture remains clear and smooth even under high-speed motion, without blurring or abrupt changes. Taking a frame of a player running quickly as an example, in 8-bit depth video, due to drastic changes in lighting, the transition between shadows and highlights on the face may appear as noticeable color blocks or breaks. However, in high-bit-depth video streams, the brightness changes in these areas are divided into more levels, resulting in a smoother and more natural gradient effect. Furthermore, when playing this video on an HDR TV, the reflected light from the grass under sunlight is more accurately reproduced, allowing viewers to see a more realistic texture of light and shadow without losing detail or distortion due to over-compression. Therefore, high-bit-depth video streams are not only a fundamental support for high-quality video experiences but also an indispensable component of modern video enhancement technologies.
[0089] In a specific embodiment, the step of performing inter-frame dynamic compensation synthesis based on the high bit-depth video data stream to obtain a quality-enhanced video output stream includes:
[0090] Dense optical flow field estimation is performed based on the high bit depth video data stream to obtain a motion vector distribution map, and motion smoothing filtering is applied to the motion vector distribution map to obtain a temporally consistent motion vector field;
[0091] Inter-frame residual prediction is performed on the high bit-depth video data stream using the time-consistent motion vector field to obtain a predicted residual map, and adaptive compensation coefficients are generated based on the predicted residual map to obtain a compensation coefficient matrix.
[0092] Based on the compensation coefficient matrix, the high bit depth video data stream is spatiotemporally fused and synthesized to obtain a synthesized video frame sequence. Based on the synthesized video frame sequence, multi-scale detail enhancement is performed to obtain a detail-enhanced video stream.
[0093] The detail-enhanced video stream is compressed using entropy coding optimization to obtain a compressed video data stream. Then, quality-adaptive filtering is performed on the compressed video data stream to obtain a quality-enhanced video output stream.
[0094] Specifically, the process of performing inter-frame dynamic compensation synthesis based on the high bit-depth video data stream to obtain the enhanced video output stream is the final integration step in the entire IPTV video quality enhancement method to achieve improved inter-frame consistency and optimized visual smoothness. This step ensures that the enhanced video not only achieves high bit-depth performance in single-frame image quality but also possesses high consistency and natural transition effects in the temporal dimension through multiple key technology modules such as dense optical flow field estimation, motion vector smoothing filtering, inter-frame residual prediction, adaptive compensation mechanism, spatiotemporal fusion synthesis, multi-scale detail enhancement, and entropy coding optimization. In specific implementation, the system first performs dense optical flow field estimation based on the high bit-depth video data stream, that is, models the pixel-level motion relationship between consecutive video frames to obtain the displacement direction and amplitude of each pixel in each frame, thereby generating a motion vector distribution map. For example, in live sports broadcasts, when cameras switch rapidly or players run at high speed, the footage contains numerous complex motion patterns. In such cases, the system employs dense optical flow algorithms like TV-L1 to calculate precise pixel motion trajectories frame by frame, generating a motion vector distribution map of 1920×1080 pixels per frame. To eliminate motion jitter caused by noise or mismatches, the system further performs motion smoothing filtering on the motion vector distribution map. This involves suppressing unstructured motion noise and preserving the true motion trend through methods such as local mean filtering or bilateral filtering, resulting in a temporally consistent motion vector field. For instance, in a football match video with frequent camera cuts, the original motion vector map might contain as many as 5% to 8% anomalous vector points. After smoothing filtering, this proportion can be reduced to below 1%, significantly improving the accuracy of subsequent inter-frame compensation. Subsequently, the system performs inter-frame residual prediction on the high bit-depth video data stream using the temporally consistent motion vector field. Specifically, based on the motion vector information of the current frame, it extracts and aligns corresponding image blocks from the reference frame, then performs pixel-level difference calculations between the aligned image and the current frame to generate a predicted residual map. This map reflects the degree of inconsistency between frames and can be used to guide subsequent compensation strategies. For example, in a certain frame, the goal area may lose some details due to motion blur. In this case, the residual value of that area in the predicted residual map will reach a high level (e.g., the average residual value exceeds 30 gray levels), indicating that stronger compensation is needed. Based on this, the system generates adaptive compensation coefficients based on the predicted residual map, automatically adjusting the compensation weights according to the residual intensity of different regions to generate a compensation coefficient matrix. For example, for regions with residual values greater than 20, the system assigns a compensation coefficient of 1.5; while for regions with residual values less than 10, it assigns only a compensation coefficient of 1.0, ensuring that while improving image quality, it avoids artifacts caused by over-enhancement.Next, the system performs spatiotemporal fusion synthesis on the high-bit-depth video data stream based on the compensation coefficient matrix. This involves weighted fusion of the current frame and reference frame using the compensation coefficients to generate a synthesized video frame sequence. This process fully considers the spatial structural similarity and temporal continuity between frames, ensuring that the enhanced video frames retain their original details while further improving frame consistency and transition smoothness. For example, in a scene of an athlete quickly jumping to catch a ball, the original video might exhibit slight jumping or edge blurring. However, after spatiotemporal fusion synthesis, the system can preserve the sense of speed while making the figure's outline clearer and the boundary transition between the background and foreground smoother. Subsequently, the system performs multi-scale detail enhancement based on the synthesized video frame sequence. This involves using techniques such as Laplacian pyramids or wavelet transforms to enhance high-frequency details of the image at different scale levels, such as texture, edges, and contrast, thereby generating a detail-enhanced video stream. For example, in a scene of a densely packed audience, the system enhances facial contours and expression details at high-frequency levels while strengthening overall illumination uniformity at low-frequency levels, making the image more layered and realistic. Finally, the system performs data compression processing on the enhanced video stream using entropy coding optimization. This involves employing efficient coding strategies (such as HEVC or AV1) to control the bitrate and optimize the bitstream of the enhanced video data, generating a compressed video data stream. During this process, the system dynamically adjusts the QP value (quantization parameter) based on the content complexity of each frame. For example, QP=24 is set for high-texture areas, while QP=28 is set for low-texture areas, thus minimizing bitrate overhead while maintaining image quality. Subsequently, the system performs adaptive quality filtering on the compressed video data stream. This involves deblocking, loop filtering, or adaptive sharpening of the reconstructed image at the decoding end to eliminate blockiness and blurring caused by compression, ultimately outputting a quality-enhanced video output stream. For example, when playing this optimized video on an HDR display device, the green of the grass on the field is more vibrant, the blue of the sky is more transparent, and the details of the players' movements are clearly visible even in shadow areas, resulting in a significant improvement in overall image quality. In summary, this step, through multiple key stages including dense optical flow estimation, motion vector smoothing, inter-frame residual prediction, adaptive compensation, spatiotemporal fusion, multi-scale enhancement, entropy coding optimization, and quality filtering, achieves comprehensive optimization and high-quality synthesis of high bit-depth video data streams. In the entire IPTV video quality enhancement process, this step not only effectively solves problems such as inter-frame discontinuity, blurred details, and compression distortion inherent in traditional video transmission, but also ensures high temporal consistency and visual comfort in the enhanced video by introducing a dynamic compensation mechanism. Especially in typical application scenarios with frequent dynamic scenes and high detail requirements, such as live sports broadcasts, this inter-frame dynamic compensation synthesis mechanism demonstrates stronger robustness and adaptability, thereby significantly improving the overall video quality and user viewing experience.
[0095] In a specific embodiment, the step of performing spatiotemporal fusion synthesis on the high bit-depth video data stream based on the compensation coefficient matrix to obtain a synthesized video frame sequence includes:
[0096] Multi-scale spatial feature extraction is performed on the high bit-depth video data stream to obtain a spatial gradient feature map, and image segmentation is performed based on the spatial gradient feature map to obtain a video content region mask.
[0097] The video content region mask is adaptively weighted using the compensation coefficient matrix to obtain a dynamic fusion weight map, and the high bit depth video data stream is pixel-level compensated based on the dynamic fusion weight map to obtain a compensated pixel mapping table.
[0098] Based on the compensated pixel mapping table, the high bit-depth video data stream is subjected to spatiotemporal domain fusion processing to obtain a spatiotemporally aligned video frame sequence, and edge-aware filtering is performed on the spatiotemporally aligned video frame sequence to obtain a smoothly transitioned video frame sequence.
[0099] The smooth transition video frame sequence is reconstructed using Laplacian pyramids to perform multi-resolution synthesis, resulting in a high-resolution video frame sequence. Color consistency correction is then performed on the high-resolution video frame sequence to obtain a synthesized video frame sequence.
[0100] Specifically, the process of spatiotemporal fusion synthesis of the high bit-depth video data stream based on the compensation coefficient matrix to obtain a synthetic video frame sequence is one of the core steps in the entire IPTV video quality enhancement method to achieve coordinated optimization of spatial detail and temporal consistency. This step involves multiple technical steps, including multi-scale spatial feature extraction, image content segmentation, dynamic weight calculation, pixel-level compensation, spatiotemporal alignment, edge-aware filtering, Laplacian pyramid reconstruction, and color consistency correction, ultimately generating a synthetic video frame sequence that meets high-quality standards in both spatial resolution and temporal coherence. In the specific implementation process, the system first performs multi-scale spatial feature extraction on the high bit-depth video data stream, that is, using a Gaussian-Laplacian hybrid filter bank or the Canny edge detection algorithm to extract the spatial gradient information of each frame at different scales, thereby generating a spatial gradient feature map. For example, in the application scenario of live sports events, high-frequency structures such as stripes on players' jerseys and the arrangement of seats in the stands will show obvious edge response values (such as gradient magnitudes exceeding 40 gray levels) in the spatial gradient feature map, while the background area will show low gradient values (such as less than 10 gray levels). This process provides a basis for subsequent image content segmentation. Subsequently, the system performs image segmentation based on the spatial gradient feature map, dividing the image into multiple semantically meaningful video content regions and generating video content region masks. For example, in a football match scene, the system divides the image into multiple regions such as "athlete subject," "field surface," "spectator stands," and "sky," assigning a unique identifier to each region to form a binary mask image of size 1920×1080. This region segmentation method helps to apply differentiated processing strategies based on the visual importance and motion characteristics of different regions in subsequent processing. Next, the system performs adaptive weight calculation on the video content region mask using the compensation coefficient matrix, that is, dynamically adjusting the weight distribution of each region in the fusion process by combining the texture complexity, motion intensity, and visual sensitivity of each region, thereby generating a dynamic fusion weight map. For example, in a certain frame, the "athlete's face" region contains rich detail information and is in a high-speed motion state, so the system assigns it a higher fusion weight (e.g., 0.85), while for static and simple texture regions such as the "sky," only a lower weight (e.g., 0.3) is assigned. This ensures that critical areas receive stronger enhancements during spatiotemporal fusion, while non-critical areas are treated moderately to conserve resources. Furthermore, the system performs pixel-level compensation on the high bit-depth video data stream based on the dynamic fusion weight map. Specifically, it corrects each pixel point-by-point based on the difference between the current frame and the reference frame, combined with the quantization compensation value in the compensation coefficient matrix, generating a compensated pixel mapping table. For example, in a fast-paced ball-passing shot, motion blur causes brightness attenuation at the edges of the sphere. The system applies a +15 to +20 grayscale compensation value to these edge pixels, making the sphere's outline sharper and clearer.This mapping table records the compensated output value of each pixel in each frame, providing accurate data support for subsequent spatiotemporal fusion. Based on this, the system performs spatiotemporal fusion processing on the high-bit-depth video data stream using the compensated pixel mapping table. Specifically, it combines the inter-frame motion vector field with the spatial compensation results, weightedly fusing the information from the current frame with the preceding and following frames to generate a spatiotemporally aligned video frame sequence. For example, in a continuously played 50-frame video, the system employs a bidirectional prediction mechanism, extracting corresponding pixels from the (n-1)th and (n+1)th frames respectively and performing a weighted average. The weighting coefficients are dynamically adjusted based on the smoothness of the motion vector and the compensation intensity, ensuring a natural and smooth transition between frames and avoiding jumps or flickering. Subsequently, the system performs edge-aware filtering processing on the spatiotemporally aligned video frame sequence, using nonlocal mean filtering or anisotropic diffusion algorithms to smooth edge regions in the image while preserving their structural integrity, thereby generating a smoothly transitioning video frame sequence. For example, during a player's run, the high speed may cause jagged edges or jitter in some frames. In this case, the system applies stronger filtering coefficients to these areas (e.g., expanding the filtering window to 7×7) to make the edge transitions smoother without affecting the overall continuity of the movement. Finally, the system performs multi-resolution synthesis on the smooth transition video frame sequence using Laplacian pyramid reconstruction. This involves decomposing the image into multiple frequency levels, enhancing details at each level, and then fusing them to generate a high-resolution video frame sequence. For example, in a close-up shot, minute details such as sweat beads and muscle textures on a player's face are usually distributed in the high-frequency levels. The system enhances these levels to make them more clearly visible in the final high-definition output. Afterward, the system performs color consistency correction based on the high-resolution video frame sequence. This involves matching and normalizing the color histograms of adjacent frames to eliminate color shifts caused by compensation processing or encoding compression, ultimately outputting a synthesized video frame sequence. In summary, this step achieves refined spatiotemporal fusion processing of high bit-depth video data streams through multiple key technologies, including multi-scale spatial feature extraction, content region mask generation, dynamic fusion weight calculation, pixel-level compensation, spatiotemporal fusion, edge-aware filtering, Laplacian pyramid reconstruction, and color consistency correction. In the entire IPTV video quality enhancement process, this step not only effectively solves common problems in traditional fusion methods such as edge blurring, inter-frame jumps, and color inconsistencies, but also ensures that the enhanced video meets high-quality standards in both spatial resolution and temporal continuity by introducing a dynamic compensation mechanism. Especially in typical application scenarios with frequent dynamic scenes and high detail requirements, such as live sports broadcasts, this spatiotemporal fusion synthesis mechanism demonstrates stronger structure preservation capabilities and visual comfort, thereby significantly improving the overall performance of video quality and the user viewing experience.
[0101] The above describes an IPTV video quality enhancement method according to an embodiment of the present invention. The following describes an IPTV video quality enhancement system according to an embodiment of the present invention. Please refer to [link / reference]. Figure 2 One embodiment of an IPTV video quality enhancement system according to the present invention includes:
[0102] Separation module 21 is used to perform inter-frame separation on IPTV video streams to obtain video frame sequences and inter-frame temporal characteristics;
[0103] Mapping module 22 is used to perform adaptive color gamut mapping on the video frame sequence based on the inter-frame temporal features to obtain color reconstruction sequence data;
[0104] Recombination module 23 is used to perform tensor decomposition and recombination on the color reconstruction sequence data to obtain a multi-dimensional texture feature map;
[0105] Optimization module 24 is used to perform dynamic bit depth optimization on the multi-dimensional texture feature map to obtain a high bit depth video data stream;
[0106] The synthesis module 25 is used to perform inter-frame dynamic compensation synthesis based on the high bit depth video data stream to obtain a quality-enhanced video output stream.
[0107] In this embodiment, the specific implementation of each module in the above system embodiment is described in the above method embodiment, and will not be repeated here.
[0108] Reference Figure 3 This invention also provides a computer device whose internal structure can be as follows: Figure 3 As shown, the computer device includes a processor, memory, display screen, input device, network interface, and database connected via a system bus. The processor provides computing and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores the operating system, computer programs, and the database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The database stores the data corresponding to this embodiment. The network interface is used for communication with external terminals via a network connection. When the computer program is executed by the processor, it implements the above-described method.
[0109] Those skilled in the art will understand that Figure 3 The structures shown are merely block diagrams of some structures related to the present invention and do not constitute a limitation on the computer devices on which the present invention is applied.
[0110] An embodiment of the present invention also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the above-described method. It is understood that the computer-readable storage medium in this embodiment can be a volatile readable storage medium or a non-volatile readable storage medium.
[0111] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the present invention and embodiments can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual-rate SDRAM (SSRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM, etc.
[0112] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, apparatus, article, or method that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, apparatus, article, or method. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, apparatus, article, or method that includes that element.
[0113] The above description is only a preferred embodiment of the present invention and does not limit the patent scope of the present invention. Any equivalent structural or procedural transformations made based on the content of the present invention specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of the present invention.
Claims
1. A method for enhancing IPTV video quality, characterized in that, Includes the following steps: Inter-frame separation is performed on the IPTV video stream to obtain the video frame sequence and inter-frame temporal characteristics; Based on the inter-frame temporal features, the video frame sequence is subjected to adaptive color gamut mapping to obtain color reconstruction sequence data; wherein, the color reconstruction sequence data includes chromaticity component reconstruction coefficients for describing the enhancement ratio of each chromaticity channel, luminance mapping curves for controlling the luminance enhancement intensity, and a set of color gamut calibration parameters for adapting to the color spaces of different display devices. Tensor decomposition and recombination are performed on the color reconstruction sequence data to obtain a multi-dimensional texture feature map; wherein, the multi-dimensional texture feature map includes a texture edge gradient map for describing the changes in edge intensity in the image, with higher values indicating clearer edges, a structural saliency map for identifying regions of visual attention in the image, and a texture orientation field intensity distribution map for characterizing the directional consistency of the texture. Dynamic bit-depth optimization is performed on the multi-dimensional texture feature map to obtain a high bit-depth video data stream; Based on the high bit depth video data stream, inter-frame dynamic compensation synthesis is performed to obtain a quality-enhanced video output stream; The step of dynamically optimizing the multi-dimensional texture feature map to obtain an optimized video data stream includes: Bit depth quantization analysis is performed on the multi-dimensional texture feature map to obtain a quantization error distribution map, and local bit depth statistics are performed based on the quantization error distribution map to obtain a bit depth distribution histogram. The bit depth distribution histogram is segmented and quantized using a non-uniform quantizer to obtain a multi-level quantization parameter set. Local bit allocation is then performed on the multi-level quantization parameter set to obtain a bit allocation weight table. Based on the bit allocation weight table, adaptive quantization compensation is performed on the texture region in the multi-dimensional texture feature map to obtain the texture compensation coefficient matrix, and bit depth optimization reconstruction is performed on the texture compensation coefficient matrix to obtain the bit depth reconstruction mapping table. The video stream is reassembled by dynamically mapping the bit depth to obtain a high bit depth video data stream. The high bit depth video data stream includes a quantization compensation coefficient describing the error correction magnitude of each pixel, a bit depth optimization parameter indicating the final bit depth configuration of each region, and a bit depth mapping curve defining the mapping relationship between different gray levels.
2. The IPTV video quality enhancement method according to claim 1, characterized in that, The process of performing inter-frame separation on the IPTV video stream to obtain the video frame sequence and inter-frame temporal characteristics includes: The IPTV video stream is subjected to temporal signal sampling analysis to obtain a video sampling data stream, and the video sampling data stream is subjected to frame synchronization mark extraction to obtain a frame synchronization feature matrix. Frame boundary localization is performed on the frame synchronization feature matrix to obtain a video frame segmentation sequence, and inter-frame motion vector calculation is performed based on the video frame segmentation sequence to obtain inter-frame displacement feature vectors; Based on the inter-frame displacement feature vector, the video frame segmentation sequence is temporally correlated and grouped to obtain a video frame sequence, and temporal features are extracted from the video frame sequence to obtain inter-frame temporal features.
3. The IPTV video quality enhancement method according to claim 1, characterized in that, The step of performing adaptive color gamut mapping on the video frame sequence based on the inter-frame temporal features to obtain color reconstruction sequence data includes: Color component decoupling analysis is performed on the inter-frame temporal features to obtain the RGB three-channel component mapping matrix, and color space coordinate transformation is performed based on the RGB three-channel component mapping matrix to obtain the YUV color space feature parameters. The YUV color gamut feature parameters are nonlinearly stretched by chromaticity histogram equalization to obtain the color distribution equalization coefficient, and chromaticity saturation is corrected based on the color distribution equalization coefficient to obtain the chromaticity correction mapping table. The video frame sequence is remapped for luminance channel based on the chroma correction mapping table to obtain a luminance enhancement sequence, and then chroma channel compensation is performed on the luminance enhancement sequence to obtain a chroma compensation vector. The chromaticity compensation vector is dynamically mapped by chromaticity gamma calibration to obtain color reconstruction sequence data.
4. The IPTV video quality enhancement method according to claim 1, characterized in that, The process of performing tensor decomposition and recombination on the color reconstruction sequence data to obtain a multi-dimensional texture feature map includes: The color reconstruction sequence data is constructed using a third-order tensor to obtain a spatiotemporal color feature tensor, and the nuclear norm minimization decomposition is performed based on the spatiotemporal color feature tensor to obtain the tensor core coefficient matrix. The tensor core coefficient matrix is subjected to multimodal separation by Tucker decomposition to obtain three subspace projection matrices in the time domain, spatial domain, and color domain. Feature correlation analysis is then performed on the three subspace projection matrices to obtain a cross-modal feature correlation map. Texture features are extracted based on the cross-modal feature association map to obtain multi-scale texture descriptors, and the multi-scale texture descriptors are spatially reorganized to obtain a texture structure reconstruction matrix. The texture structure reconstruction matrix is dimensionally fused using singular value constraints to obtain a multi-dimensional texture feature map.
5. The IPTV video quality enhancement method according to claim 1, characterized in that, The process of performing inter-frame dynamic compensation synthesis based on the high bit-depth video data stream to obtain a quality-enhanced video output stream includes: Dense optical flow field estimation is performed based on the high bit depth video data stream to obtain a motion vector distribution map, and motion smoothing filtering is applied to the motion vector distribution map to obtain a temporally consistent motion vector field; Inter-frame residual prediction is performed on the high bit-depth video data stream using the time-consistent motion vector field to obtain a predicted residual map, and adaptive compensation coefficients are generated based on the predicted residual map to obtain a compensation coefficient matrix. Based on the compensation coefficient matrix, the high bit depth video data stream is spatiotemporally fused and synthesized to obtain a synthesized video frame sequence. Based on the synthesized video frame sequence, multi-scale detail enhancement is performed to obtain a detail-enhanced video stream. The detail-enhanced video stream is compressed using entropy coding optimization to obtain a compressed video data stream. Then, quality-adaptive filtering is performed on the compressed video data stream to obtain a quality-enhanced video output stream.
6. The IPTV video quality enhancement method according to claim 5, characterized in that, The process of performing spatiotemporal fusion synthesis on the high bit-depth video data stream based on the compensation coefficient matrix to obtain a synthesized video frame sequence includes: Multi-scale spatial feature extraction is performed on the high bit-depth video data stream to obtain a spatial gradient feature map, and image segmentation is performed based on the spatial gradient feature map to obtain a video content region mask. The video content region mask is adaptively weighted using the compensation coefficient matrix to obtain a dynamic fusion weight map, and the high bit depth video data stream is pixel-level compensated based on the dynamic fusion weight map to obtain a compensated pixel mapping table. Based on the compensated pixel mapping table, the high bit-depth video data stream is subjected to spatiotemporal domain fusion processing to obtain a spatiotemporally aligned video frame sequence, and edge-aware filtering is performed on the spatiotemporally aligned video frame sequence to obtain a smoothly transitioned video frame sequence. The smooth transition video frame sequence is reconstructed using Laplacian pyramids to perform multi-resolution synthesis, resulting in a high-resolution video frame sequence. Color consistency correction is then performed on the high-resolution video frame sequence to obtain a synthesized video frame sequence.
7. An IPTV video quality enhancement system, characterized in that, include: The separation module is used to perform inter-frame separation of IPTV video streams to obtain video frame sequences and inter-frame temporal characteristics; The mapping module is used to perform adaptive color gamut mapping on the video frame sequence based on the inter-frame temporal features to obtain color reconstruction sequence data; wherein, the color reconstruction sequence data includes chromaticity component reconstruction coefficients for describing the enhancement ratio of each chromaticity channel, luminance mapping curves for controlling the luminance enhancement intensity, and a set of color gamut calibration parameters for adapting to the color spaces of different display devices. The recombination module is used to perform tensor decomposition and recombination on the color reconstruction sequence data to obtain a multi-dimensional texture feature map. The multi-dimensional texture feature map includes a texture edge gradient map that describes the changes in edge intensity in the image, with higher values indicating clearer edges; a structural saliency map that identifies the regions of visual attention in the image; and a texture orientation field intensity distribution map that characterizes the directional consistency of the texture. The optimization module is used to perform dynamic bit-depth optimization on the multi-dimensional texture feature map to obtain a high bit-depth video data stream; The synthesis module is used to perform inter-frame dynamic compensation synthesis based on the high bit depth video data stream to obtain a quality-enhanced video output stream. The step of dynamically optimizing the multi-dimensional texture feature map to obtain an optimized video data stream includes: Bit depth quantization analysis is performed on the multi-dimensional texture feature map to obtain a quantization error distribution map, and local bit depth statistics are performed based on the quantization error distribution map to obtain a bit depth distribution histogram. The bit depth distribution histogram is segmented and quantized using a non-uniform quantizer to obtain a multi-level quantization parameter set. Local bit allocation is then performed on the multi-level quantization parameter set to obtain a bit allocation weight table. Based on the bit allocation weight table, adaptive quantization compensation is performed on the texture region in the multi-dimensional texture feature map to obtain the texture compensation coefficient matrix, and bit depth optimization reconstruction is performed on the texture compensation coefficient matrix to obtain the bit depth reconstruction mapping table. The video stream is reassembled by dynamically mapping the bit depth to obtain a high bit depth video data stream. The high bit depth video data stream includes a quantization compensation coefficient describing the error correction magnitude of each pixel, a bit depth optimization parameter indicating the final bit depth configuration of each region, and a bit depth mapping curve defining the mapping relationship between different gray levels.
8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Video stream image adaptive enhancement method and system
CN115511755A
Method for improving detail quality of high-dynamic infrared digital image based on weighted gamma correction
CN119168905A