Video image compression algorithm based on H.266 / VVC standard
By combining adaptive quantization strategy and asymmetric quantization matrix with generative adversarial network, the blocking effect and edge preservation problems in the H.266/VVC standard video image compression algorithm are solved, and the quality and compression efficiency of video images are improved.
Patent Information
- Application Number
- CN202510903209.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-01
- Publication Date
- 2025-09-19
AI Technical Summary
The existing video image compression algorithm based on the H.266/VVC standard has significant blocking effects in complex texture areas, quantization distortion in areas sensitive to the human eye, and insufficient edge preservation capabilities of traditional filters, resulting in poor video image quality and compression efficiency.
An adaptive quantization strategy is introduced to generate a visual attention heat map through a spatiotemporal saliency detection model, adjust the QP value of the coding tree unit, and design an asymmetric quantization matrix to optimize the transformation quantization effect. At the same time, a generative adversarial network is used to enhance the subjective visual quality of the reconstructed frame.
It significantly improves the subjective visual quality and compression efficiency of video images, meets the needs of modern high-definition video transmission, reduces the blocking effect in complex texture areas and the quantization distortion in areas sensitive to the human eye, and optimizes edge preservation capabilities.
Smart Images

Figure CN120676160A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of video coding and image processing, and in particular to a video image compression algorithm based on the H.266 / VVC standard. Background Art
[0002] In the application of modern video compression technology, the H.266 / VVC (Versatile Video Coding) standard has gradually become a hot research topic in the industry due to its higher compression efficiency and wide adaptability. However, in practical applications, existing video image compression algorithms based on the H.266 / VVC standard still face many challenges, especially in terms of significant blocking artifacts in complex texture areas, quantization distortion in areas sensitive to the human eye, and the insufficient edge preservation ability of traditional filters. Further optimization is urgently needed to meet the requirements of high-quality video transmission.
[0003] After searching, a method for image compression and decompression with the publication number CN101563926B was found. This patent defines pixel clusters and dynamically selects clustering methods according to the image content, and calculates a control parameter group for post-processing operations for each cluster to improve the quality of the decompressed image. However, this technical solution fails to fully consider the block effect problem when processing complex texture areas, which may result in obvious visual artifacts in high-detail areas. In addition, its quantization strategy for areas sensitive to the human eye lacks an adaptive mechanism, which may affect the subjective visual experience at low bit rates. At the same time, the solution relies on traditional filters for post-processing, with limited edge preservation capabilities and certain limitations in the recovery of high-frequency information, thereby reducing the clarity of the reconstructed image.
[0004] In addition, a method for decompressing image data is disclosed with the publication number CN112422980B. This patent uses multi-level bit encoding to achieve efficient image data decompression by reading the channel origin value and difference representation. Although this solution can improve compression efficiency and reduce storage requirements, its processing ability for complex texture areas is weak, and block effects are prone to occur at high compression ratios, affecting the overall image quality. At the same time, this method does not design a special quantization strategy for areas sensitive to the human eye, which may lead to the loss or distortion of key information. In addition, because it adopts a fixed-mode decoding process, it lacks flexibility and is difficult to cope with diverse needs in different scenarios, especially in terms of edge sharpness maintenance, resulting in blurred edges of the reconstructed image, affecting the visual experience.
[0005] The above-mentioned problems indicate that existing video image compression algorithms based on the H.266 / VVC standard still have significant deficiencies in terms of suppressing blocking artifacts in complex texture regions, controlling quantization distortion in areas sensitive to the human eye, and maintaining edge preservation. Therefore, the present invention provides a video image compression algorithm based on the H.266 / VVC standard. This algorithm aims to significantly improve the subjective visual quality and compression efficiency of video images by introducing an adaptive quantization strategy, enhancing the blocking artifact suppression mechanism in complex texture regions, and optimizing the design of edge-preserving filters, thereby meeting the requirements of modern high-definition video transmission. Summary of the Invention
[0006] The present invention aims to address the shortcomings of existing technologies by providing a video image compression algorithm based on the H.266 / VVC standard. This algorithm addresses the problems of significant blocking artifacts in complex textured areas, quantization distortion in areas sensitive to the human eye, and the inadequate edge-preserving capabilities of traditional filters. By introducing an adaptive quantization strategy, enhancing blocking artifact suppression mechanisms in complex textured areas, and optimizing edge-preserving filter design, the algorithm improves the subjective visual quality and compression efficiency of video images.
[0007] In order to solve the above technical problems, the following technical solutions are adopted: A first aspect of the present invention provides a video image compression algorithm based on the H.266 / VVC standard, comprising the following steps: Step A: Generate a visual attention heat map through the spatiotemporal saliency detection model.
[0008] First, a spatiotemporal saliency detection model is constructed. This model integrates static saliency and dynamic motion features, calculates the SAD (Sum of Absolute Differences) between adjacent frames to generate a motion intensity map, and then weightedly fuses the motion intensity map with the static saliency map to form the final visual attention heat map. The static saliency map is generated through local contrast analysis, while the motion intensity map is generated by summing the absolute differences between pixel values in adjacent frames.
[0009] Step B: Adjust the QP value of the coding tree unit according to the visual attention heat map partition.
[0010] Based on the visual attention heat map generated in step A, the coding tree units are partitioned, with each partition corresponding to a saliency level. Subsequently, the quantization parameter (QP value) of each coding tree unit is adjusted based on the saliency level. Higher saliency levels are assigned lower QP values to reduce quantization distortion in high-attention areas. Specifically, the QP value adjustment range is determined by a preset QP offset table, which is pre-set based on experimental data to ensure quantization distortion control at different saliency levels.
[0011] Step C: In the transform and quantization stage, a first asymmetric quantization matrix is used for the intra-frame prediction block, and a second asymmetric quantization matrix is used for the inter-frame prediction block.
[0012] During the transform quantization process, a first asymmetric quantization matrix is designed for the intra-frame prediction block. This matrix reduces the quantization error of the high-frequency components by increasing the quantization step size of the high-frequency components. A second asymmetric quantization matrix is designed for the inter-frame prediction block. This matrix improves the quantization accuracy of the low-frequency components by reducing the quantization step size of the low-frequency components. The selection of the two quantization matrices is based on the input of the prediction mode signal to ensure the optimal quantization effect under different prediction modes.
[0013] Step D: In the loop filtering stage, the reconstructed frame is input into the pre-trained generative adversarial network for enhancement processing.
[0014] The loop filtering stage receives the reconstructed frame as input and feeds it into a pre-trained generative adversarial network for enhancement processing. The generative adversarial network consists of a generator and a discriminator. The generator contains a 7-level residual dense block, which is used to extract multi-scale features and generate enhanced images. The discriminator uses a Markov discriminator structure to determine the similarity between the generated image and the original image. The training process of the generative adversarial network uses the Vimeo-90K dataset to construct training sample pairs to ensure that the network has good generalization ability.
[0015] As a further improvement, in step A, the spatiotemporal saliency detection model includes a static saliency map generation module and a motion intensity map generation module. The static saliency map generation module is used to generate a static saliency map; The motion intensity map generation module is used to generate a motion intensity map, and the static saliency map and the motion intensity map are weightedly fused to generate a visual attention heat map.
[0016] Specifically, the method for generating a static saliency map is: First, a local contrast analysis is performed on the current frame to calculate the brightness difference between each pixel and its neighboring pixels to form a local contrast map. Then, the local contrast map is Gaussian smoothed to remove noise interference. Finally, the smoothed local contrast map is normalized to obtain a static saliency map.
[0017] As a further improvement, in step A, the method for generating the motion intensity map is: First, the brightness difference of corresponding pixels in adjacent frames is calculated, and its absolute value is taken to form a preliminary motion map; then the preliminary motion map is Gaussian filtered to smooth the motion boundaries; finally, the smoothed motion map is normalized to obtain a motion intensity map.
[0018] As a further improvement, in step B, the visual attention heat map is divided into multiple partitions, each partition corresponds to a significance level, the higher the significance level, the lower the assigned quantization parameter, and the quantization parameter adjustment range is determined by a preset quantization parameter offset table, which is pre-set based on experimental data.
[0019] Specifically, the design method of the QP offset table is: Through experiments, subjective visual quality score data at different saliency levels was collected to establish a mapping relationship between saliency level and QP offset value. Specifically, the saliency level was divided into five intervals, each corresponding to a fixed QP offset value. The QP offset value range was [-6, 6] to ensure the fineness of quantization parameter adjustment.
[0020] As a further improvement, in step C, the design method of the first asymmetric quantization matrix is: According to the frequency characteristics of the intra-frame prediction block, the quantization step size is increased by 10% for high-frequency components and reduced by 10% for low-frequency components; at the same time, the quantization step size in the diagonal direction remains unchanged to ensure the preservation of image edge information.
[0021] As a further improvement, in step C, the design method of the second asymmetric quantization matrix is: Based on the frequency characteristics of the inter-frame prediction block, the quantization step size is reduced by 15% for low-frequency components and increased by 5% for high-frequency components. At the same time, the quantization step sizes in the horizontal and vertical directions are kept consistent to ensure quantization accuracy in flat areas of the image.
[0022] As a further improvement, in step D, the specific structure of the 7-level residual dense block of the generator is: Each level of residual dense block contains three convolutional layers with convolution kernel sizes of 3×3, 5×5 and 7×7, respectively, for extracting multi-scale features; each level of residual dense blocks is connected by skip connections to ensure that shallow features can be passed to the deep network; the output of the last level of residual dense block passes through the global average pooling layer and the fully connected layer to generate an enhanced image.
[0023] As a further improvement, in step D, the specific design of the Markov discriminator structure of the discriminator is: The discriminator consists of 5 convolutional layers. The convolution kernel size of the first 4 convolutional layers is 4×4, with a stride of 2, which is used to gradually reduce the size of the feature map; the convolution kernel size of the 5th convolutional layer is 3×3, with a stride of 1, which is used to extract global features; the output of the discriminator is a single-channel feature map, which represents the local similarity between the generated image and the original image.
[0024] A second aspect of the present invention provides a corresponding device, which includes a processor and a memory storing program instructions. The processor is configured to execute the aforementioned video image compression algorithm based on the H.266 / VVC standard when running the program instructions.
[0025] A third aspect of the present invention provides a corresponding storage medium, which stores program instructions. When the program instructions are run, they execute the aforementioned video image compression algorithm based on the H.266 / VVC standard.
[0026] The video image compression algorithm based on the H.266 / VVC standard provided by the embodiment of the present disclosure combines the spatiotemporal saliency detection model to generate a visual attention heat map, implements an adaptive quantization strategy by adjusting the QP value of the coding tree unit through partitioning, designs an asymmetric quantization matrix to optimize the transformation quantization effect, and uses a generative adversarial network to enhance the subjective visual quality of the reconstructed frame. The above effectively solves the problems of significant block effects in complex texture areas, quantization distortion in areas sensitive to the human eye, and insufficient edge preservation capabilities of traditional filters, significantly improving the subjective visual quality and compression efficiency of video images and meeting the needs of modern high-definition video transmission. Specifically, it has the following technical effects: Saliency Detection Model: This paper introduces a spatiotemporal saliency detection model to generate a visual attention heatmap and adjust quantization parameters based on this. This approach allocates resources based on human visual characteristics, prioritizing details in important areas, thereby improving subjective quality.
[0027] Asymmetric quantization matrix: Different asymmetric quantization matrices are designed for intra-frame prediction blocks and inter-frame prediction blocks. This fine-tuning method can better adapt to different types of image content and improve compression efficiency.
[0028] Generative Adversarial Network Enhancement: A pre-trained generative adversarial network (GAN) is used to enhance the reconstructed frames during the loop filtering stage. This is a technique rarely used in traditional compression algorithms and helps to further improve the quality of the reconstructed image. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] The present invention will be further described below in conjunction with the accompanying drawings: Figure 1 This is a schematic diagram of the overall process of a video image compression algorithm based on the H.266 / VVC standard provided by an embodiment of the present invention.
[0030] Figure 2 Schematic diagram of the structure of the spatiotemporal domain saliency detection model in an embodiment of the present invention.
[0031] Figure 3 A schematic diagram of the design of an asymmetric quantization matrix in an embodiment of the present invention. Figure 4Schematic diagram of the structure of a generative adversarial network in an embodiment of the present invention.
[0032] The accompanying drawings are numbered as follows: 1. Static saliency map generation module; 2. Motion intensity map generation module; 3. Visual attention heat map; 4. First asymmetric quantization matrix; 5. Second asymmetric quantization matrix; 6. Generator; 7. Discriminator. DETAILED DESCRIPTION
[0033] To make the objectives, technical solutions, and advantages of the present invention more clearly understood, the present invention is described in detail below with reference to the accompanying drawings and examples. However, it should be understood that the specific embodiments described herein are merely illustrative of the present invention and are not intended to limit the scope of the present invention. In addition, descriptions of well-known structures and technologies are omitted in the following description to avoid unnecessary confusion of the present invention.
[0034] The present invention provides a video image compression algorithm based on the H.266 / VVC standard, and its specific implementation is as follows. Figure 1 The embodiment of the present invention provides a schematic diagram of the overall process of the video image compression algorithm based on the H.266 / VVC standard. The core of the present invention is to combine the saliency detection model, the asymmetric quantization matrix and the generative adversarial network to form a complete video compression framework. Compared with the method in the prior art that simply relies on fixed rules or empirical formulas, the present invention proposes a more intelligent and adaptive solution. Specifically, the present invention shows the main steps from spatiotemporal saliency detection to loop filtering enhancement processing. The algorithm effectively improves the subjective visual quality and compression efficiency of video images by introducing an adaptive quantization strategy, a complex texture area block effect suppression mechanism, and an optimized edge-preserving filter design.
[0035] In the specific implementation process, step A is first performed, that is, generating visual attention heat through the spatiotemporal saliency detection model. Figure 3 . This process relies on the collaborative work of the static saliency map generation module 1 and the motion intensity map generation module 2. The static saliency map generation module 1 calculates the brightness difference between each pixel and its neighboring pixels by performing local contrast analysis on the current frame to form a local contrast map. The local contrast map is then Gaussian smoothed to remove noise interference, and the smoothed local contrast map is normalized to obtain a static saliency map. The motion intensity map generation module 2 is responsible for calculating the brightness difference between corresponding pixels in adjacent frames, taking its absolute value to form a preliminary motion map, and performing Gaussian filtering on the preliminary motion map to smooth the motion boundary. Finally, the smoothed motion map is normalized to obtain a motion intensity map. The static saliency map and the motion intensity map are combined by weighted fusion to form the final visual attention heat map. Figure 3The specific implementation of weighted fusion is to perform linear weighted summation of the static saliency map and the motion intensity map according to preset weights. The weight values are pre-set based on experimental data to ensure a reasonable balance between the two saliency features.
[0036] Next, execute step B, based on the visual attention heat Figure 3 Partition the coding tree unit and adjust the quantization parameter (QP value). Figure 3 It is divided into multiple partitions, each of which corresponds to a significance level. The higher the significance level, the lower the assigned QP value, thereby reducing the quantization distortion in high-attention areas. The adjustment range of the QP value is determined by the preset QP offset table. The design method of the QP offset table is to collect subjective visual quality rating data under different significance levels through experiments, and establish a mapping relationship between the significance level and the QP offset value. Specifically, the significance level is divided into 5 intervals, each of which corresponds to a fixed QP offset value. The range of the QP offset value is [-6, 6], which ensures the fineness of the adjustment of the quantization parameter. The core of this process is to use the visual attention thermal Figure 3 The spatial distribution characteristics of guide the dynamic adjustment of QP value, thereby realizing the adaptive quantization strategy.
[0037] Then, step C is executed, where different asymmetric quantization matrices are used for the intra-frame prediction block and the inter-frame prediction block during the transform quantization stage. The first asymmetric quantization matrix 4 is used for the intra-frame prediction block. Its design method is to increase the quantization step size by 10% for high-frequency components and reduce it by 10% for low-frequency components based on the frequency characteristics of the intra-frame prediction block. At the same time, the quantization step size in the diagonal direction remains unchanged to ensure the preservation of image edge information. The second asymmetric quantization matrix 5 is used for the inter-frame prediction block. Its design method is to reduce the quantization step size by 15% for low-frequency components and increase it by 5% for high-frequency components based on the frequency characteristics of the inter-frame prediction block. The quantization step size in the horizontal and vertical directions is kept consistent to ensure quantization accuracy in flat areas of the image. The selection of the two quantization matrices is based on the input of the prediction mode signal to ensure the optimal quantization effect under different prediction modes. The key to this process is to design the corresponding quantization matrix according to the frequency characteristics of different prediction blocks, thereby optimizing the transform quantization effect.
[0038] Then, step D is executed. In the loop filtering stage, the reconstructed frame is input into the pre-trained generative adversarial network for enhancement processing. The generative adversarial network consists of a generator 6 and a discriminator 7. Its structure is as follows: Figure 4As shown in the figure, generator 6 consists of seven levels of residual dense blocks. Each level of residual dense blocks contains three convolutional layers with kernel sizes of 3×3, 5×5, and 7×7, respectively, for extracting multi-scale features. Each level of residual dense blocks is connected by skip connections to ensure that shallow-level features can be passed to the deeper layers of the network. The output of the last level of residual dense blocks passes through a global average pooling layer and a fully connected layer to generate an enhanced image. Discriminator 7 adopts a Markov discriminator structure and consists of five convolutional layers. The convolution kernel size of the first four convolutional layers is 4×4 with a stride of 2, which is used to gradually reduce the size of the feature map. The convolution kernel size of the fifth convolutional layer is 3×3 with a stride of 1, which is used to extract global features. The output of discriminator 7 is a single-channel feature map, which represents the local similarity between the generated image and the original image. The training process of the generative adversarial network uses the Vimeo-90K dataset to construct training sample pairs to ensure the network has good generalization ability. The focus of this process is to use the generator 6 and discriminator 7 of the generative adversarial network to work together to enhance the reconstructed frames to improve the subjective visual quality.
[0039] The connection and position relationship of each module and component in the above steps are as follows. Static saliency map generation module 1 and motion intensity map generation module 2 together constitute the spatiotemporal saliency detection model, which generates visual attention heat through weighted fusion. Figure 3 Visual attention heat Figure 3 This is passed as an input signal to the QP value adjustment module in step B, guiding the partitioning of the coding tree units and the dynamic adjustment of the QP value. The first asymmetric quantization matrix 4 and the second asymmetric quantization matrix 5 are applied to the transform quantization stage of the intra-frame prediction block and the inter-frame prediction block, respectively. Their selection is controlled by the prediction mode signal. The generator 6 and the discriminator 7 form a generative adversarial network, which receives the reconstructed frame from the loop filtering stage as input and generates an enhanced image output. The modules and components in the entire algorithm work closely together to complete the video image compression task.
[0040] In practical application scenarios, the video image compression algorithm proposed in this invention is applicable to the field of high-definition video transmission and storage. For example, in video conferencing systems, this algorithm can effectively reduce the blocking effect in complex texture areas, improve the subjective visual quality in areas sensitive to the human eye, and at the same time reduce the video bit rate to meet real-time transmission requirements. In video surveillance systems, this algorithm can optimize the design of edge-preserving filters and enhance the detail of reconstructed frames, thereby improving the clarity and recognition of monitoring images. In addition, in video-on-demand platforms, this algorithm can optimize compression efficiency through adaptive quantization strategies and asymmetric quantization matrices, reducing storage costs and improving user experience.
[0041] The technical solution of the present invention generates visual attention heat through the spatiotemporal saliency detection model Figure 3, combined with an adaptive quantization strategy to adjust the QP value, designing an asymmetric quantization matrix to optimize the transform quantization effect, and utilizing a generative adversarial network to enhance the subjective visual quality of the reconstructed frames. These technical approaches effectively address the issues of significant blocking artifacts in complex texture areas, quantization distortion in areas sensitive to the human eye, and the insufficient edge preservation capabilities of traditional filters. This significantly improves the subjective visual quality and compression efficiency of video images, meeting the requirements of modern high-definition video transmission.
[0042] In order to better enable relevant personnel in this technical field to fully understand and implement the present invention, the specific implementation principle of the present invention is further supplemented below with reference to a specific application scenario.
[0043] In the actual application of the video conferencing system, step A is first performed to generate the visual attention heat map through the collaborative work of the static saliency map generation module 1 and the motion intensity map generation module 2. Figure 3 The static saliency map generation module 1 calculates the brightness difference between each pixel and its neighboring pixels by performing local contrast analysis on the current frame to form a local contrast map, and removes noise interference through Gaussian smoothing and normalizes it to obtain a static saliency map. The motion intensity map generation module 2 calculates the brightness difference between the corresponding pixels of adjacent frames, takes its absolute value to form a preliminary motion map, and then smoothes the motion boundary through Gaussian filtering and normalizes it to obtain a motion intensity map. Subsequently, the static saliency map and the motion intensity map are linearly weighted and summed according to the preset weights to generate the final visual attention heat map. Figure 3 During this process, the weight values are pre-set based on experimental data to ensure a reasonable balance between the two saliency features, thereby accurately reflecting the areas of human eye focus on the video content.
[0044] Next, execute step B, based on the visual attention heat Figure 3 Partition the coding tree unit and adjust the quantization parameter (QP value). Figure 3 It is divided into multiple partitions, each of which corresponds to a significance level. The higher the significance level, the lower the QP value is assigned to reduce the quantization distortion of the high attention area. The adjustment range of the QP value is determined by the preset QP offset table, which is established by experimentally collecting subjective visual quality score data under different significance levels. Specifically, the significance level is divided into 5 intervals, each of which corresponds to a fixed QP offset value in the range of [-6, 6]. The core of this process is to use the thermal model of visual attention to determine the QP value. Figure 3 The spatial distribution of the QP value guides the dynamic adjustment of the QP value, thus realizing an adaptive quantization strategy. For example, in a video conferencing scenario, the face area usually has a higher saliency level, so a lower QP value is assigned to preserve more detail information and improve subjective visual quality.
[0045] Then, step C is performed, where different asymmetric quantization matrices are used for intra-frame prediction blocks and inter-frame prediction blocks during the transform and quantization phase. The first asymmetric quantization matrix 4 is used for intra-frame prediction blocks. Based on the frequency characteristics of intra-frame prediction blocks, the quantization step size is increased by 10% for high-frequency components and decreased by 10% for low-frequency components. Meanwhile, the diagonal quantization step size remains unchanged to ensure the preservation of image edge information. The second asymmetric quantization matrix 5 is used for inter-frame prediction blocks. Based on the frequency characteristics of inter-frame prediction blocks, the quantization step size is reduced by 15% for low-frequency components and increased by 5% for high-frequency components. The horizontal and vertical quantization step sizes are kept consistent to ensure quantization accuracy in flat areas of the image. The selection of the two quantization matrices is based on the prediction mode input signal to ensure optimal quantization performance for different prediction modes. For example, in video conferencing scenarios, the background area is often an inter-frame prediction block. Using the second asymmetric quantization matrix 5 can reduce quantization error in low-frequency components, thereby improving the smoothness and consistency of the background area.
[0046] Then, step D is executed. In the loop filtering stage, the reconstructed frame is input into the pre-trained generative adversarial network for enhancement processing. The generative adversarial network consists of a generator 6 and a discriminator 7. Its structure is as follows: Figure 4 As shown in the figure, generator 6 consists of seven levels of residual dense blocks. Each level of residual dense blocks contains three convolutional layers with kernel sizes of 3×3, 5×5, and 7×7, respectively, for extracting multi-scale features. Each level of residual dense blocks is connected by skip connections to ensure that shallow-level features can be passed to the deeper layers of the network. The output of the last level of residual dense blocks passes through a global average pooling layer and a fully connected layer to generate an enhanced image. Discriminator 7 adopts a Markov discriminator structure and consists of five convolutional layers. The convolution kernel size of the first four convolutional layers is 4×4 with a stride of 2, which is used to gradually reduce the size of the feature map. The convolution kernel size of the fifth convolutional layer is 3×3 with a stride of 1, which is used to extract global features. The output of discriminator 7 is a single-channel feature map, which represents the local similarity between the generated image and the original image. The training process of the generative adversarial network uses the Vimeo-90K dataset to construct training sample pairs to ensure the network has good generalization ability. For example, in video conferencing scenarios, generative adversarial networks can effectively enhance the detail information in reconstructed frames, reduce the blocking effect in complex texture areas, and improve the clarity and edge sharpness of facial areas.
[0047] The connection and position relationship of each module and component in the above steps are as follows. Static saliency map generation module 1 and motion intensity map generation module 2 together constitute the spatiotemporal saliency detection model, which generates visual attention heat through weighted fusion. Figure 3 Visual attention heat Figure 3This is passed as an input signal to the QP value adjustment module in step B, guiding the partitioning of the coding tree units and the dynamic adjustment of the QP value. The first asymmetric quantization matrix 4 and the second asymmetric quantization matrix 5 are applied to the transform quantization stage of the intra-frame prediction block and the inter-frame prediction block, respectively. Their selection is controlled by the prediction mode signal. The generator 6 and the discriminator 7 form a generative adversarial network, which receives the reconstructed frame from the loop filtering stage as input and generates an enhanced image output. The modules and components in the entire algorithm work closely together to complete the video image compression task.
[0048] In practical applications of video surveillance systems, this algorithm can optimize the design of edge-preserving filters and enhance the detail of reconstructed frames, thereby improving the clarity and recognizability of surveillance images. For example, in nighttime surveillance scenarios, the combined application of the first asymmetric quantization matrix 4 and the second asymmetric quantization matrix 5 can effectively reduce quantization distortion in low-light conditions. Simultaneously, the enhancement processing of the generative adversarial network can enhance the edge sharpness and detail clarity of target objects, thereby improving the recognizability of surveillance images.
[0049] In addition, in the actual application of video on demand platform, this algorithm can optimize compression efficiency through adaptive quantization strategy and asymmetric quantization matrix, reduce storage cost and improve user experience. Figure 3 Guiding the dynamic adjustment of QP values can effectively reduce bit rate usage. At the same time, the enhanced processing of the generative adversarial network can improve the overall viewing quality of the video and meet users' demand for high-definition picture quality.
[0050] This method significantly improves the subjective quality of video while maintaining a high compression rate by adjusting quantization parameters using partitions, employing asymmetric quantization matrices, and performing GAN enhancement. Experimental results show that compared to traditional H.266 / VVC encoders, this algorithm achieves an average PSNR improvement of approximately 0.5dB at the same bit rate, and significantly improves the SSIM value.
[0051] Cost and efficiency: Despite the introduction of a saliency detection model and GAN, these modules can be deployed through pre-training, thus not significantly increasing the computational burden at runtime. Furthermore, through the refined management of quantization parameters, this invention effectively reduces storage and transmission costs.
[0052] Any content not described in detail in the specification belongs to the prior art known to those skilled in the art, and the model parameters of each electrical appliance are not specifically limited, and conventional equipment can be used. In this technical solution, electrical control components not mentioned are not shown in the figures because they belong to the prior art and will not be described here.
[0053] The above are only specific embodiments of the present invention, but the technical features of the present invention are not limited thereto. Any simple changes, equivalent substitutions, or modifications based on the present invention to solve substantially the same technical problems and achieve substantially the same technical effects are included within the scope of protection of the present invention.
Claims
1. A video image compression algorithm based on the H.266 / VVC standard, characterized in that: The following steps are involved: Step A: Generate a visual attention heat map using a spatiotemporal saliency detection model (3); Step B, partitioning and adjusting the quantization parameters of the coding tree units according to the visual attention heat map (3); Step C: In the transform quantization stage, a first asymmetric quantization matrix (4) is used for the intra-frame prediction block, and a second asymmetric quantization matrix (5) is used for the inter-frame prediction block; Step D: In the loop filtering stage, the reconstructed frame is input into the pre-trained generative adversarial network for enhancement processing. The generative adversarial network includes a generator (6) and a discriminator (7).
2. The video image compression algorithm based on the H.266 / VVC standard according to claim 1, characterized in that: In step A, the spatiotemporal saliency detection model includes a static saliency map generation module (1) and a motion intensity map generation module (2). The static saliency map generation module (1) is used to generate a static saliency map; The motion intensity map generation module (2) is used to generate a motion intensity map, and the static saliency map and the motion intensity map are weightedly fused to generate the visual attention heat map (3).
3. The video image compression algorithm based on the H.266 / VVC standard according to claim 2, characterized in that: The static saliency map generation module (1) calculates the brightness difference between each pixel and its neighboring pixels by performing local contrast analysis on the current frame to form a local contrast map, and performs Gaussian smoothing and normalization on the local contrast map to generate a static saliency map.
4. The video image compression algorithm based on the H.266 / VVC standard according to claim 2, characterized in that: The motion intensity map generation module (2) generates a preliminary motion map by calculating the absolute value of the brightness difference between corresponding pixel points in adjacent frames, and performs Gaussian filtering and normalization on the preliminary motion map to generate a motion intensity map.
5. The video image compression algorithm based on the H.266 / VVC standard according to claim 1, characterized in that: In step B, the visual attention heat map (3) is divided into a plurality of partitions, each partition corresponding to a significance level. The higher the significance level, the lower the quantization parameter assigned. The adjustment range of the quantization parameter is determined by a preset quantization parameter offset table, which is pre-set according to experimental data.
6. The video image compression algorithm based on the H.266 / VVC standard according to claim 1, characterized in that: In step C, the first asymmetric quantization matrix (4) is used for intra-frame prediction blocks, the quantization step size of the high-frequency components of the first asymmetric quantization matrix (4) is increased by 10%, the quantization step size of the low-frequency components is reduced by 10%, and the quantization step size in the diagonal direction remains unchanged; the second asymmetric quantization matrix (5) is used for inter-frame prediction blocks, the quantization step size of the low-frequency components of the second asymmetric quantization matrix (5) is reduced by 15%, the quantization step size of the high-frequency components is increased by 5%, and the quantization step sizes in the horizontal and vertical directions remain consistent.
7. The video image compression algorithm based on the H.266 / VVC standard according to claim 1, characterized in that: In step D, the generator (6) includes seven levels of residual dense blocks, each level of residual dense blocks includes three convolution layers, and the convolution kernel sizes are three times three, five times five, and seven times seven, respectively. Each level of residual dense blocks is connected by jump connections, and the output of the last level of residual dense blocks is subjected to a global average pooling layer and a fully connected layer to generate an enhanced image.
8. The video image compression algorithm based on the H.266 / VVC standard according to claim 1, characterized in that: In step D, the discriminator (7) includes five convolutional layers, the convolution kernel size of the first four convolutional layers is four times four, the step size is two, the convolution kernel size of the fifth convolutional layer is three times three, the step size is one, and the output of the discriminator (7) is a single-channel feature map.
9. A device, characterized in that: The system comprises a processor and a memory storing program instructions, wherein the processor is configured to execute the video image compression algorithm based on the H.266 / VVC standard as claimed in any one of claims 1 to 8 when running the program instructions.
10. A storage medium, characterized in that: Program instructions are stored, and when the program instructions are run, the video image compression algorithm based on the H.266 / VVC standard according to any one of claims 1 to 8 is executed.
Citation Information
Patent Citations
Image compression and decompression
CN101563926B
Image data decompression
CN112422980B
Cited By
Ultra-high-definition video stream adaptive coding method based on deep learning visual saliency
CN121397231A