Image processing method, video processing system, device, medium and program product

The Lightweight Learning Loop Filter (LLILF) solves the problem of high computational complexity in neural network loop filtering tools, achieving efficient and real-time image filtering processing to meet the encoding requirements of high-definition video.

CN121099074APending Publication Date: 2025-12-09MIGU CO LTD +3
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511169229.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-20
Publication Date
2025-12-09

AI Technical Summary

Technical Problem

Existing loop filtering tools are based on neural network models with complex structures and a huge number of parameters, resulting in high computational complexity. This makes them difficult to deploy efficiently in real-time encoding scenarios and cannot meet the real-time filtering requirements of high-resolution videos.

Method used

A lightweight learning loop filter (LLILF) is adopted, which consists of a cascaded first convolutional unit, a nonlinear processing unit, and a second convolutional unit. The first model is trained through knowledge distillation, which reduces the number of parameters and computational complexity and improves filtering efficiency.

Benefits of technology

It fulfills the requirement for real-time filtering in high-resolution video encoding, improving image processing efficiency and picture quality while reducing the dependence on computing resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121099074A_ABST
    Figure CN121099074A_ABST
Patent Text Reader

Abstract

The invention discloses an image processing method, a video processing system, equipment, a medium and a program product. Comprises: acquiring a first image; performing filtering processing on the first image through a loop filter to obtain a second image; wherein the loop filter comprises a first model, and the first model comprises a first convolution unit, a nonlinear processing unit and a second convolution unit. Through the scheme, real-time filtering can be realized.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of video coding, and particularly relates to an image processing method, a video processing system, a device, a medium and a program product. BACKGROUND

[0002] With the rapid development of high-definition video, ultra-high definition (UHD) video, virtual reality (VR) and streaming media services, higher requirements are put forward for the coding efficiency of video data. In the coding process, loop filtering is a key link. The loop filtering algorithm in the third generation of digital audio and video coding and decoding standard (AVS3) can only perform data processing based on fixed rules and signal characteristics, and lacks adaptive ability to complex image content. In order to solve the above technical problems, the related technology provides a technical scheme of introducing a loop filtering tool constructed by a neural network into a coding architecture.

[0003] However, the loop filtering tool constructed based on the neural network has a complex structure, resulting in a large amount of parameters and high data processing complexity, and thus leading to low efficiency of loop filtering. SUMMARY

[0004] In order to solve the above technical problems, the embodiments of the present application provide an image processing method, a video processing system, a device, a medium and a program product.

[0005] The embodiments of the present application first provide an image processing method, comprising:

[0006] obtaining a first image;

[0007] performing filtering processing on the first image by a loop filter to obtain a second image;

[0008] The loop filter comprises a first model, and the first model comprises a first convolution unit, a nonlinear processing unit and a second convolution unit.

[0009] The embodiments of the present application also provide a video processing system, comprising:

[0010] an input unit configured to obtain a video;

[0011] a processing unit configured to process the video to obtain a first image;

[0012] a loop filter configured to filter the first image to obtain a second image, wherein the loop filter comprises a first model comprising a first convolution unit, a non-linear processing unit, and a second convolution unit;

[0013] an encoding unit configured to encode the second image to obtain a data stream.

[0014] The embodiments of the present application further provide an electronic device, which comprises a processor and a memory; the memory stores a computer program; and the computer program is executable by the processor to implement the image processing method.

[0015] The embodiments of the present application further provide a computer readable storage medium, which comprises a computer program; and the computer program is executable by the processor to implement the image processing method.

[0016] The embodiments of the present application further provide a computer program product, which comprises a computer program; and the computer program is executable by the processor to implement the image processing method.

[0017] The image processing method provided by the embodiments of the present application comprises a loop filter, the loop filter comprises a first model, the first model comprises a first convolution unit, a non-linear processing unit and a second convolution unit, so it can be known that the first model provided by the embodiments of the present application has a simple structure, and accordingly, the first model has a small number of parameters, thereby the data processing efficiency of the first model can be improved; on this basis, the filtering efficiency of the loop filter can be improved; at the same time, after the first image is obtained, the first image is filtered by the loop filter to obtain a second image, thereby the efficiency of filtering the first image by the loop filter can be improved, and then the real-time filtering requirement of the first image can be met. BRIEF DESCRIPTION OF DRAWINGS

[0018] Figure 1 a flowchart of the image processing method provided by the embodiments of the present application;

[0019] Figure 2 a structure diagram of the teacher model provided by the embodiments of the present application;

[0020] Figure 3 a structure diagram of the ResBlock provided by the embodiments of the present application;

[0021] Figure 4 a structure diagram of the first model provided by the embodiments of the present application;

[0022] Figure 5 a structure diagram of the first model training provided by the embodiments of the present application;

[0023] Figure 6 A structural diagram of a video coding framework for deploying a first model is provided for an embodiment of the present application;

[0024] Figure 7 A structural diagram of a video processing system is provided for an embodiment of the present application;

[0025] Figure 8 A structural diagram of an electronic device is provided for an embodiment of the present application. DETAILED DESCRIPTION

[0026] The technical solutions in the embodiments of the present application will be clearly and completely described in connection with the drawings in the embodiments of the present application.

[0027] It should be understood that the specific embodiments described herein are merely intended to explain the present application and are not intended to limit the present application.

[0028] With the rapid development of high-definition video, UHD, VR, and streaming services, the generation and consumption of video data are growing exponentially, which drives the continuous evolution of video compression and coding technology. In the field of video coding, AVS3 adopts a more complex coding framework and multiple coding tools to improve compression efficiency.

[0029] Among the above coding tools, as a key component, the loop filtering module plays an important role in improving the quality of compressed images. This module usually includes units such as Deblocking Filter (DBF), Sample Adaptive Offset (SAO), and Adaptive Loop Filtering (ALF) based on region classification.

[0030] Although the loop filtering algorithm on which the traditional loop filtering module in AVS3 realizes its functions is designed ingeniously, the difficulty of optimizing this loop filtering algorithm increases as the demand for improving coding efficiency increases. Moreover, the filters in the above loop filtering algorithm process signal features based on fixed rules, lacking the ability to adapt to complex image content, making it difficult to adapt to image filtering needs in different scenarios.

[0031] To solve the above technical problems, with the development of deep learning technology, loop filtering modules based on neural networks have been gradually introduced into the coding framework. These new loop filtering modules have shown advantages in improving the subjective perception quality and objective evaluation indicators of compressed video, making them one of the research hotspots in loop filtering technology.

[0032] Meanwhile, in the research progress of the AVS3 standard, the neural network-based hybrid video coding technology is also deepening, and the related research is actively exploring to embed the coding tools realized through the neural network into the AVS3 reference software to realize the organic integration of the standard tool chain and the intelligent method.

[0033] The coding tools include neural network loop filtering tools, which take a convolutional neural network (CNN) structure as the main framework, introduce channel attention or spatial attention mechanisms to enhance the modeling capability of local features by stacking convolutional blocks, and learn the distribution characteristics of artifacts to effectively enhance image quality. However, this method introduces high network calculation complexity while improving the filtering effect, making it difficult to be efficiently deployed in real-time encoding scenarios.

[0034] In related technologies, neural networks are also applied to replace the deblocking effect filter and sample adaptive offset module in the Versatile Video Coding (VVC) standard. This method extracts features based on the reconstructed information at the Coding Tree Unit (CTU) level, the partition depth matrix, the Slice type, and the Quantization Parameter (QP) value to learn the distortion characteristics and output the code stream flag, thereby shortening the overall encoding time under the condition of equivalent encoding gain. However, although this scheme introduces some prior information to reduce network redundancy, its network structure is still relatively complex, and the processing of prior information is limited to simple matrix multiplication and splicing operations, and the lightweight potential has not been fully tapped.

[0035] Moreover, the neural network-based loop filtering tools in related technologies use deep CNN structures that include a large number of residual blocks, which can easily result in a per-pixel computation of the loop filtering tools reaching hundreds or even thousands of multiply-accumulate operations (kMACs / pixel); correspondingly, the parameter size of the complex deep CNN structure is usually in the hundreds of thousands to millions, which not only increases the parameter storage cost, but also puts higher requirements on the inference framework, device memory, and bandwidth, and is contrary to the requirements of the video coding standard for module lightweight and standardization.

[0036] In summary, the neural network-based loop filtering tools in related technologies have complex model structures and large parameter quantities, resulting in high computational complexity; correspondingly, the above models are difficult to achieve real-time inference on the CPU in the encoding process of high-resolution videos (such as 1080p / 4K), thereby failing to meet the real-time requirements of image filtering.

[0037] To solve the above technical problems, embodiments of the present application provide an image processing method, a video processing system, a device, a medium and a program product.

[0038] Figure 1 A flowchart of the image processing method provided by the embodiments of the present application is shown in Figure 1 The method can include the following steps:

[0039] Step 101, obtaining a first image.

[0040] In an embodiment, the first image can include any image frame in a video to be encoded, and can also include the block data obtained after the image frame is blocked.

[0041] In an embodiment, the first image can be obtained by the following way:

[0042] Any frame image in a video is blocked to obtain block data, an image prediction result is obtained through inter-frame prediction and intra-frame prediction processing in a video coding and decoding framework, a first residual data between the block data and the image prediction result is calculated, a first processing result is obtained by transforming / quantizing the first residual data, a second processing result is obtained by inverse transforming / inverse quantizing the first processing result, a second residual data between the second processing result and the image prediction result is calculated, and then the second residual data is determined as the first image.

[0043] Step 102, filtering the first image through a loop filter to obtain a second image.

[0044] The loop filter includes a first model; the first model includes a first convolution unit, a nonlinear processing unit and a second convolution unit.

[0045] In an embodiment, the first model can include a neural network model, which can improve the picture quality of the first image; illustratively, the first model can include a neural network loop filter (NNLF).

[0046] In an embodiment, illustratively, the first convolution unit can include M first convolution kernels, where M can be 8, the size of each first convolution kernel can be the same or different, and if the size of each first convolution kernel is the same, the size of the first convolution kernel can be 3x3.

[0047] In an embodiment, the nonlinear processing unit can include a nonlinear function, such as ReLU.

[0048] In an embodiment, the second convolution unit can include N second convolution kernels, where N can be the same as or different from M, and N can also be 8, and each second convolution kernel can have the same or different size, and if each second convolution kernel has the same size, the size of each second convolution kernel can be 3*3.

[0049] It should be noted that M and N are both integers greater than 1.

[0050] In an embodiment, the parameters of the first convolution unit and the second convolution unit can be different.

[0051] In an embodiment, the arrangement among the first convolution unit, the second convolution unit, and the nonlinear processing unit in the first model can be adjusted or determined according to the processing requirements of the in-loop filter on the first image.

[0052] In an embodiment, the first model can have 153 parameters; for example, because the first model has a simple structure and a small number of parameters, the first model can be referred to as a lightweight learning-based in-loop filter (LLILF).

[0053] In an embodiment, in addition to including the first model, the in-loop filter can also include a DBF, an SAO, and an ALF; for example, the DBF, the first model, the SAO, and the ALF can be sequentially cascaded to form the in-loop filter.

[0054] Correspondingly, the filtering processing of the first image by the in-loop filter to obtain the second image can be implemented in the following manner:

[0055] The DBF is used to perform deblocking processing on the first image to obtain first data; the first model is used to filter the first data to obtain second data; the SAO is used to perform ringing removal and pixel value optimization on the second data to obtain third data; and the ALF is used to process the third data to obtain the second image.

[0056] In an embodiment, because the first model includes the first convolution unit, the nonlinear processing unit, and the second convolution unit, the first model can be used to efficiently process the first data, thereby improving the coding gain and reducing the coding time.

[0057] Specifically, the first convolution unit can extract features of a first dimension in the first data to obtain a first feature extraction result; the nonlinear processing unit can extract nonlinear features in the first feature extraction result to capture nonlinear structure information in the first image; and the second convolution unit can extract features of a second dimension in the nonlinear features to obtain a second feature extraction result; where the first dimension and the second dimension can be different.

[0058] As can be seen from the above, the image processing method provided by the embodiments of the present application has a loop filter including a first model, and the first model includes a first convolution unit, a nonlinear processing unit, and a second convolution unit. Therefore, the first model provided by the embodiments of the present application has a simple structure, and accordingly, the first model has a small number of parameters, thereby improving the data processing efficiency of the first model. On this basis, the filtering efficiency of the loop filter can be improved. At the same time, after obtaining the first image, the loop filter is used to filter the first image to obtain a second image, thereby improving the efficiency of filtering the first image by the loop filter, and further meeting the demand for real-time filtering of the first image.

[0059] Based on the foregoing embodiments, in the image processing method provided by the embodiments of the present application, the first model is distilled from the second model; and the first model is composed of the first convolution unit, the nonlinear processing unit, and the second convolution unit arranged in cascade.

[0060] In an implementation manner, the second model can be a teacher model; accordingly, the first model can be a trained student model; and exemplarily, the trained student model can be deployed in the loop filter.

[0061] In an implementation manner, the teacher model can include a deep neural network model; exemplarily, the teacher model can include a third convolution unit, a set of residual learning units, and a fourth convolution unit arranged in cascade; where the number of residual learning units in the set of residual learning units can be greater than or equal to X; X can be an integer greater than 1; and exemplarily, the value of X can be 20.

[0062] Exemplarily, by stacking X residual learning units in the teacher model, the accuracy of learning high-quality filtering expression of the teacher model can be improved.

[0063] Figure 2 A structural diagram of the teacher model provided by the embodiments of the present application is shown in FIG. 1. Figure 2As shown, the teacher model 200 can include a third convolution unit, a first ReLU, a residual unit set, a fourth convolution unit, and a second ReLU; wherein the third convolution unit and the fourth convolution unit can respectively include 8 3x3 convolution kernels, the third convolution unit is used to extract basic features, and the fourth convolution unit is used to fuse and reconstruct the basic features to output an enhanced brightness image, and the residual learning unit in the residual set unit can be a ResBlock, and the structure thereof is as shown in Figure 3

[0064] Figure 3 The structure diagram of the ResBlock provided by the embodiments of the present application is shown in Figure 3 As shown, each ResBlock 300 can be composed of a 3x3 convolution kernel, a ReLU, and a 3x3 convolution kernel arranged in cascade; for example, Figure 2 The yth ResBlock in the ResBlock set 300 can obtain the output data of the previous data processing unit, perform convolution processing on the output data through the Conv 3x3 convolution kernel to obtain a third processing result, then perform nonlinear processing on the third processing result through the ReLU included therein to obtain a fourth processing result, perform convolution processing on the fourth processing result through the second Conv 3x3 convolution kernel to obtain a fifth processing result, and finally perform residual calculation on the output data and the fifth processing result to obtain the output result thereof; wherein y can be an integer greater than or equal to 1, if the value of y is 1, the previous data processing unit can include the first ReLU, and if the value of y is greater than 1, the previous data processing unit can include the y-1th ResBlock.

[0065] In an embodiment, the teacher model can be obtained by pre-training based on sample data; for example, the sample data can include at least part of the B, C, and D class data in the Big Video Dataset for Deep Video Compression (BVI-DVC).

[0066] For example, the teacher model can be trained in the following manner:

[0067] The reconstructed frame image in the video coding framework is input to the initial state of the teacher model, and the predicted image is obtained through the processing of the teacher model, and then the mean square error between the original image and the predicted image is used to optimize the parameters of the initial state of the teacher model, so as to realize the training of the teacher model; wherein the mean square error can be calculated by formula (1):

[0068]

[0069] wherein I is the reconstructed frame image for the original image, f(I; θ​t G is the predicted image. q The image is the original image, Q is the number of samples, and θ is the number of samples. t These are the parameters of the teacher model.

[0070] Figure 4 A schematic diagram of the structure of the first model provided in the embodiments of this application, as shown below. Figure 4 As shown, the first model 400 consists of a first convolutional unit, a nonlinear processing unit, and a second convolutional unit arranged in cascades; wherein the first convolutional unit and the second convolutional unit may each include eight 3×3 convolutional kernels, and the nonlinear processing unit may include a ReLU function.

[0071] Combination Figure 2 to Figure 4 It can be seen that the structural complexity of the first model is significantly reduced compared to the teacher model. Correspondingly, the number of parameters in the first model is also significantly reduced compared to the teacher model. For example, in the embodiments of this application, the number of parameters in the first model can be 153, and its computational complexity is 0.161kMACs / pixel, thereby reducing the dependence on data processing resources.

[0072] For example, the size of the input and output data of the first model is 1×W×H.

[0073] In one implementation, the first model is obtained by distillation of the second model, which can be achieved in the following way:

[0074] The first model is obtained by learning the image filtering knowledge of the teacher model through knowledge distillation-driven learning of the student model; specifically, the above process can be implemented in the following way:

[0075] The student model extracts features from the sample data to obtain a first set; the teacher model extracts features from the sample data to obtain a second set; based on the degree of difference between the first and second sets, the parameters of the student model are adjusted. This process is then repeated recursively using both the adjusted student model and the teacher model to gradually adjust the parameters of the student model until the degree of difference between the first and second sets is less than or equal to a threshold. This parameter adjustment process yields a student model with adjusted parameters, which can then acquire all the image filtering knowledge from the teacher model. In this case, the student model with adjusted parameters can be designated as the first model.

[0076] In one implementation, image filtering knowledge may include the sum of methods, parameters, and processes for quality optimization of the reconstructed image, elimination of residual noise, and blur distortion.

[0077] From the above, in the image processing method provided by the embodiment of the application, the first model is distilled from the second model, so that, under the premise that the second model has the data processing function that can be realized by the NNLF, the ability and level of the first model in improving the picture quality of the reconstructed first image can be improved; and the first model is composed of the first convolution unit, the nonlinear processing unit and the second convolution unit arranged in cascade, so that not only the associated combination relationship between the units in the first model is clear, but also the tracking and capturing ability of the first model for complex features can be improved.

[0078] Based on the foregoing embodiment, in the image processing method provided by the embodiment of the application, the first model is trained by the first loss and the second loss.

[0079] The first loss is a pixel reconstruction loss; and the second loss is used to represent a distillation loss.

[0080] In an implementation, the pixel reconstruction loss can include the difference between the pixel points of the image generated in the student model parameter adjustment process and the pixel points of the original image in the sample data; for example, the pixel reconstruction loss can be calculated by formula (2):

[0081]

[0082] wherein, θ s is the parameter of the student model, and P is the sample quantity.

[0083] In an implementation, the distillation loss is used to represent the difference between the teacher output data processed by the teacher model for the sample data and the student output data processed by the student model for the sample data.

[0084] In an implementation, the first model is trained by the first loss and the second loss, which can be realized in the following manner:

[0085] The training loss is determined based on the first loss and the second loss; the parameter of the student model is adjusted gradually based on the training loss, the sample data is processed by the student model after the parameter adjustment, the sample data is processed by the teacher model, the first loss, the second loss and the training loss are determined again, and the process is repeated until the training loss is less than or equal to the loss threshold, at which time the student model trained, i.e., the first model, is obtained.

[0086] For example, the training loss can be calculated by formula (3):

[0087] L student = (1- a) L disstill + a L MSE (3)

[0088] wherein, L disstill is a distillation loss, and a is a loss weighting coefficient, which can be 0.7, to dominate the training process of the student model in terms of reconstruction quality.

[0089] Through the above training process, the first model can obtain the image filtering knowledge of the teacher model; and by using knowledge distillation to drive the training of the student model, the problems of unstable gradient and slow convergence of the student model when directly trained can be solved; at the same time, by using the intermediate layer semantic knowledge of the teacher model, the student model can learn more discriminative feature expression with lower parameter amount, so as to maintain high filtering performance and filtering quality on the premise of improving reasoning efficiency; at the same time, in the above training process, by introducing the distillation loss, the input and output features of the student model can be constrained to make it converge faster.

[0090] Exemplarily, the above training process can be implemented by an Adam optimizer, and the first learning rate of the teacher model can be 1x10 -4 , and the second learning rate of the student model is 3x10 -5 , and the first learning rate and the second learning rate can be adjusted as the training process proceeds.

[0091] Specifically, when training the teacher model, the first batch size (Batch Size) can be set to 64, and the training process can include 100 epochs, and the first learning rate can be decayed to one third of the original value after every 30 epochs; when training the student model, the second batch size (Batch Size) can be set to 8, and the same training process is 100 epochs, i.e. the second learning rate is decayed to one third of the original value after every 30 epochs.

[0092] Specifically, in the above training process, SVT-AVS3 can be used to compress sample data under random access (RA) configuration by setting four different QPs, and the compressed sample data can be randomly cropped into a plurality of 144x144 image blocks, and then data augmentation can be performed by random horizontal and vertical flipping; wherein, the value of QP can include 27, 32, 38 and 45.

[0093] Specifically, in the process of training the student model and the teacher model, the types of CPU and graphics processing unit (GPU) used respectively can be the same.

[0094] Exemplarily, the parameter configuration of the teacher model in the above training process can be as shown in Table 1, and the parameter configuration of the student model can be as shown in Table 2.

[0095] epoch 100 Batch Size 64 Loss function MSE Sample data B, C, and D class data in BVI-DVC Learning rate 1 x 10 -4 ]] Optimizer Adam

[0096] Table 1

[0097] epoch 100 Batch Size 8 Loss function MSE and distillation loss function Sample data B, C, and D class data in BVI-DVC Learning rate 3 x 10 -5 ]] Optimizer Adam

[0098] Table 2

[0099] In an implementation, the first model can replace a module for implementing loop filtering in the loop filtering link to deploy the first model in the loop filtering link, while other modules in the loop filtering link can remain unchanged.

[0100] In an implementation, after obtaining the first model, the first model can be embedded into an SVT-AVS3 encoder for performance evaluation. If the performance level of the first model is greater than or equal to the level threshold, the training can be stopped. If the performance level of the first model is less than the level threshold, the above training process can be repeatedly performed.

[0101] It should be noted that in actual applications, the artifacts and losses after video compression are mainly reflected in the bottom features such as local texture blurring and edge weakening, and such bottom information can be effectively recovered through the local convolution receptive field in the first model. Therefore, although the structure is simple, the first model still has basic image enhancement capability. And through the guided training of knowledge distillation, the first model can effectively inherit the deep perception ability of the teacher model, so that it can maintain the same filtering performance as the teacher model at low complexity.

[0102] In related technologies, in order to reduce complexity, there are technical solutions of directly designing shallow network structures to replace traditional filters, but these shallow network structures are prone to convergence difficulties, unstable performance, and obvious subjective quality decline in actual training process. And related technologies lack effective training paradigm for lightweight network structure, resulting in that the performance of these models in the encoder is worse than deep models or traditional filtering methods.

[0103] The training method of the student model provided in the embodiments of the present application can realize the training guidance of the student model through the feature map output by the teacher model, and combine the image reconstruction loss of the student model with the distillation loss, which can improve the stability of the student model training and improve the filtering performance of the first model, thereby overcoming the problem of difficult convergence of lightweight model training.

[0104] Meanwhile, in actual video compression scenarios, a filtering strategy that only pursues objective quality or compression rate is often insufficient to meet user experience, and the neural network filtering tool in the related art still has problems of insufficient subjective quality and poor visual perception in processing compression artifacts and maintaining texture details, especially in difficult encoding areas such as motion areas or high-contrast edges; and the training method provided in the embodiments of the present application can balance the perceptual effect and the calculation overhead.

[0105] As can be seen from the above, the image processing method provided in the embodiments of the present application, the first model is trained through the pixel reconstruction loss and the distillation loss, so that the pixel reconstruction loss can improve the accuracy of the first model in tracking and capturing pixel features in the image, and can also improve the training efficiency of the first model; and the distillation loss can improve the generalization ability of the first model and the capturing and migration ability of image filtering knowledge, and can also speed up the training efficiency of the first model.

[0106] Based on the foregoing embodiments, in the image processing method provided in the embodiments of the present application, the second loss is obtained according to the first feature map and the second feature map.

[0107] The first feature map includes a feature map output by a first layer of the first model, and the second feature map includes a feature map output by a first layer of the second model.

[0108] And / or,

[0109] The first feature map includes a feature map output by a last layer of the first model, and the second feature map includes a feature map output by a last layer of the second model.

[0110] In an implementation manner, the first layer of the first model can include a first convolution unit, and the last layer of the first model can include a second convolution unit; correspondingly, the first layer of the second model can include a third convolution unit, and the last layer of the second model can include a fourth convolution unit.

[0111] Correspondingly, the first feature map can include a feature map output by the first convolution unit and / or a feature map output by the second convolution unit; and the second feature map can include a feature map output by the third convolution unit and / or a feature map output by the fourth convolution unit.

[0112] Correspondingly, the second loss can represent at least one of the following:

[0113] A first feature difference between the feature map output by the first convolution unit and the feature map output by the third convolution unit;

[0114] A second feature difference between the feature map output by the second convolution unit and the feature map output by the fourth convolution unit.

[0115] the first feature difference and the second feature difference;

[0116] the first feature difference or the second feature difference.

[0117] Specifically, the sample data can include a training set and a validation set, and the sample data can be obtained by converting B, C and D type data in the BVI-DVC from an original MP4 format to a corresponding 10-bit YUV420 format; the training set can include 90% of the sample data, and the training set can include 10% of the sample data.

[0118] Figure 5 A structural diagram of student model training provided by an embodiment of the present application is shown in FIG. 1. Figure 5 The second loss can represent a feature difference between a feature map output by the third convolution unit and a feature map output by the first convolution unit, and / or a feature difference between a feature map output by the fourth convolution unit and a feature map output by the second convolution unit.

[0119] Exemplarily, in a case where the second loss, i.e., the distillation loss, includes the first feature difference and the second feature difference, the distillation loss L disstill The distillation loss function shown in equation (4) can be used to calculate the distillation loss L

[0120]

[0121] where P is the number of samples in the training set, F s (1) and F s (L) are the first feature map, respectively, F t (1) and F t (L) are the second feature map.

[0122] As can be seen from the above, the image processing method provided by the embodiment of the present application, the second loss is obtained according to the first feature map and the second feature map, and the first feature map includes a feature map output by the first layer of the first model, and the second feature map includes a feature map output by the first layer of the second model, and / or the first feature map includes a feature map output by the last layer of the first model, and the second feature map includes a feature map output by the last layer of the second model. In this way, in the above process, the second loss is obtained by the feature map alignment, so as to improve the learning accuracy of the first model to the second model.

[0123] Based on the foregoing embodiment, in the image processing method provided by the embodiment of the present application, before the first image is obtained, the following operation can also be performed:

[0124] processing the video to obtain the first image.

[0125] Accordingly, after the second image is obtained, the following operations can also be performed:

[0126] The second image is cached.

[0127] The second image is a reference image for inter-prediction.

[0128] In an embodiment, the loop filter can be arranged in a video coding framework; thus, after the video is received by the video coding framework, the video can be processed by a preceding data processing link of the loop filter in the video coding framework to obtain the first image; exemplary, the preceding data processing link can include a frame prediction module, a first residual unit and a transform processing unit.

[0129] In an embodiment, the second image is used to update a decoded image in a decoded picture buffer of the video coding framework to provide a reference image for inter-prediction of the video coding framework; exemplary, after the second image is obtained, the second image can be stored as a decoded image in the decoded picture buffer, so that a module or unit for implementing inter-prediction in the video coding framework can obtain the decoded image from the decoded picture buffer and use the decoded image as a reference image for inter-prediction, thereby improving the accuracy and pertinence of inter-prediction.

[0130] Figure 6 A structural schematic diagram of a video coding framework in which a first model is deployed is provided for embodiments of the present application. As shown in Figure 6 The video input unit in the video coding framework 600 can receive a video and perform block processing on any frame image in the video to obtain block data, and then input the block data to a frame prediction module and a first residual unit residual, respectively.

[0131] Exemplary, the frame prediction module can perform intra-prediction and inter-prediction operations to obtain a prediction image predictor; wherein the inter-prediction needs to be implemented with the aid of a decoded image stored in a decoded picture buffer.

[0132] Exemplarily, the first residual unit performs residual calculation on the prediction image predictor and the block data to obtain a first residual, and inputs the first residual to a transform / quantization unit for processing to obtain a first transform result, and then inputs the first transform result to an inv.transform / inv.quant unit and an entropy coding unit for processing; wherein the inv.transform / inv.quant unit processes the first transform result to obtain a second transform result, and inputs the second transform result to a second residual unit, and the entropy coding unit is configured to generate a video stream bitstream based on filter control data output by the in-loop filtering link and the first transform result; wherein the transform processing unit can include the transform / quantization unit, the inv.transform / inv.quant unit and the entropy coding unit.

[0133] Exemplarily, the second residual unit is configured to determine a second residual based on the prediction image predictor and the second transform result, and input the second residual to the in-loop filtering link; and the second residual can include the first image in the foregoing embodiments.

[0134] Exemplarily, the first image is processed by the in-loop filtering link to obtain a second image; wherein the LLILF in the in-loop filtering link can be the first model provided by the embodiments of the present application.

[0135] Exemplarily, after the in-loop filtering link generates the second image, the in-loop filtering link can send the second image to a decoded picture buffer to update the decoded image in the decoded picture buffer, and provide an image reference for inter-prediction.

[0136] As can be seen from the above, the image processing method provided by the embodiments of the present application processes a video to obtain a first image before obtaining the first image, and buffers a second image after obtaining the second image, and the second image is a reference image for inter-prediction. In this way, the above operations are integrated, that is, the video is processed to obtain the first image, the first image is processed to obtain the second image, and the second image is buffered to provide a reference image for inter-prediction, thereby realizing integrated processing of the video encoding process.

[0137] Based on the foregoing embodiments, the image processing method provided in the embodiments of the present application includes: a first convolution unit is configured to extract first features; a nonlinear processing unit is configured to process the first features to obtain second features; and a second convolution unit is configured to fuse and / or reconstruct the second features to obtain a second image with enhanced brightness.

[0138] In the embodiments, the first features include edge and / or texture information of the first image.

[0139] In an embodiment, the first convolution unit can be configured to extract visual features such as edge and / or texture information from the first data by using M convolution kernels; the nonlinear processing unit can be configured to extract nonlinear features from the first features to obtain the second features; the second convolution unit can be configured to process the second features by using N convolution kernels to obtain a first model output result, and then fuse and / or reconstruct the first data and the first model output result to obtain the second image with enhanced brightness.

[0140] In an embodiment, the second image with enhanced brightness can include at least part of the data in the first model output result, and can also include part of the data in the first data.

[0141] Specifically, the second image with enhanced brightness can be obtained by the following method:

[0142] The data component associated with the first channel in the first model output result and the data component associated with the second channel in the first data are integrated to obtain target data; for example, the data component associated with the first channel in the first model output result and the data component associated with the second channel in the first data can be correspondingly superimposed based on the coordinate position of the data in the first model output result to obtain the second image with enhanced brightness.

[0143] In the embodiments, the pixels in the first model output result and the first data include components associated with the first channel and the second channel.

[0144] In an embodiment, the first channel can be adjusted according to specific filtering requirements; for example, the first channel can include a brightness channel, i.e., a Y channel, so that the features associated with the brightness channel in the second features can be specifically fused and / or reconstructed; correspondingly, the second channel can include other channels in the second data except the first channel, for example, the second channel can include a chroma channel, i.e., a UV channel.

[0145] It should be noted that the size of the first data and the second image can both be 1×W×H.

[0146] From the above, the image processing method provided by the embodiment of the application can realize accurate tracking and extraction of edge and / or texture information by the first convolution unit extracting the edge and / or texture information of the first image, can realize targeted extraction of nonlinear features in the first features by the nonlinear processing unit processing the first features, and can improve the integrity of the features in the second image and can also realize targeted processing of the brightness features in the second features by the second convolution unit fusing and / or reconstructing the second features.

[0147] Based on the foregoing embodiment, the image processing method provided by the embodiment of the application can further perform the following operations:

[0148] Deriving model parameters of the first model, constructing a filter process corresponding to the first model based on the model parameters to obtain a target model, and deploying the target model to the loop filtering link.

[0149] In an implementation, the model parameters can be derived by the following way:

[0150] Deriving the model parameters in a preset data format; wherein the preset data format can include array, structure, and associated container formats; and the preset data format can be flexibly adjusted.

[0151] In an implementation, the target model can be obtained by the following way:

[0152] Obtaining a data processing link included in the filter process of the first model, and writing target code for realizing the above data processing link based on the model parameters by a programming language, and determining the target code as the target model.

[0153] In an implementation, the data transmission interface in the target model can be packaged by a first data transmission interface between DBF and NNLF in the loop filtering link and its corresponding first parameters, and a second data transmission interface between NNLF and SAO and its corresponding second parameters, and the packaged data transmission interface can be connected with the first data transmission interface and the second data transmission interface, so as to deploy the target model to the loop filtering link; wherein the first parameters can include the number and type of input and output data of the first data transmission interface, and the second parameters can include the number and type of input and output data of the second data transmission interface.

[0154] As can be seen from the above, the image processing method provided in the embodiments of the present application exports the model parameters of the first model, and constructs a filtering process corresponding to the first model based on the model parameters to obtain a target model. In this way, the same filtering function as the first model can be realized by the target model and the same filtering performance can be maintained, and the dependence on the first model is reduced. Furthermore, the target model is deployed in the loop filtering link, which can improve the flexibility of the deployment of the target model and reduce the restriction of the deployment platform and third-party factors such as the neural network on the deployment of the loop filtering link.

[0155] Based on the foregoing embodiments, in the image processing method provided in the embodiments of the present application, the model parameters of the first model can be exported in the following manner:

[0156] The model parameters are exported in a target format of a target language.

[0157] In an implementation manner, the target language can include a computer programming language. For example, the computer programming language can include an object-oriented programming language or a procedural programming language. For example, the object-oriented programming language can include C++, JAVA or C#, and the procedural programming language can include C or assembly language.

[0158] In an implementation manner, the target format can include a data format that can be read or recognized by the target language. For example, in the case that the target language is C++, the target format can include a multi-dimensional array. For example, for the first convolution unit and the second convolution unit each including 8 3x3 convolution kernels, the weight of each convolution kernel can be stored in a four-dimensional array with a dimension of 8x1x3x3, and the bias parameter corresponding to the convolution kernel can be stored in a one-dimensional array.

[0159] It should be noted that the target language can be determined or adjusted according to the deployment environment of the video coding framework device. For example, the deployment environment of the video coding framework can include the operating system and the processor of the electronic device used for deploying the video coding framework.

[0160] Correspondingly, the filtering process of the first model can be constructed based on the model parameters to obtain the target model in the following manner:

[0161] Based on the model parameters, the filtering process is constructed by the target language by using a parallel strategy to obtain the target model.

[0162] The parallel strategy includes an image block parallel processing strategy and / or an instruction parallel execution strategy.

[0163] In an implementation manner, the image block parallel processing strategy can include synchronously performing filtering processing on at least two image blocks at the same time. For example, the image block parallel processing strategy can be realized by using a process scheduling manner.

[0164] In an embodiment, the instruction parallel execution strategy can be implemented through a single instruction multiple data (SIMD) instruction set of the electronic device, and the efficiency of matrix calculation, code running and data access can be improved through the SIMD instruction set; for example, the instruction parallel execution strategy needs to rely on the function support of the processor of the electronic device to the SIMD to be implemented.

[0165] In an embodiment, the target model can be constructed in the following way:

[0166] Based on the parallel strategy, the parallel coding idea is determined, then the filtering processing flow including convolution processing contained in the filtering process is constructed based on the coding idea through the target language, and the model parameters are configured in the filtering processing flow, the encoded data is obtained, the compiled processing is performed on the encoded data to obtain the executable code, and then the executable code is determined as the target model.

[0167] For example, after obtaining the target model, the standard test sequence in the AVS common test conditions (CTC) can be used as the test set, and the video multi-method fusion quality assessment index (VMAF) is used as the evaluation standard to evaluate the performance of the proposed filter under the RA configuration; the above key test conditions are shown in Table 3:

[0168] Framework Proposed Inference Library Complexity 0.161 kMACs / pixel Number of models 1 model for luminance filtering

[0169] Table 3

[0170] In the above test process, the CPU used can be the same as the CPU used in the training stage in the foregoing embodiments.

[0171] Through the above deployment process, the target model corresponding to the LLILF can be seamlessly integrated into an open source coding framework such as SVT-AVS3, as an enhancement module between DBF and SAO, which can realize the optimization of filtering performance and engineering efficiency.

[0172] For example, after deploying the target model to the video coding framework, in the video coding framework coding process, the loop filtering link is set to all open filtering by default for each CTU of each frame image.

[0173] It can be seen from the above that the image processing method provided in the embodiments of the present application exports the model parameters in a target data format of a target language, thereby realizing flexible export of the model parameters. Furthermore, based on the model parameters, a filtering process is constructed by using the target language through a parallel strategy, and a target model is obtained. The parallel strategy includes an image block parallel processing strategy and / or an instruction parallel execution strategy. In this way, the filtering efficiency of the target model can be improved. At the same time, the two filtering processes are constructed by using the target language, which can reduce the dependence of the model on a software environment and / or a hardware environment for deploying a video coding framework, thereby improving the flexibility of the video coding framework deployment.

[0174] Specifically, in the RA configuration condition, compared with SVT-AVS3 M11, the video coding framework provided in the present application can improve the encoding quality of 1080p and 720p, and the average improvement is 3.7 VMAF scores; in the test sequence dimension; in the RA configuration condition, compared with SVT-AVS3 M11, the video coding framework provided in the present application realizes a VMAF saving of 15.35% of the average BD-rate (Bjontegaard-Delta); on the other hand, in the ultrafast configuration condition of x265 and x264, the video coding framework provided in the present application saves 9.47% and 62.65% of the BD-rate (VMAF), respectively; in the real-time dimension, the video coding framework provided in the present application can realize efficient encoding and decoding processing for 1080p@30fps and 720p@60fps; and for the high resolution support dimension, the video coding framework provided in the embodiments of the present application can still support 10fps operation under the condition of 4K; further, the model parameters of the first model provided in the embodiments of the present application are significantly reduced, thereby providing data support for the encoding standard modularization, parameter fixation and convenient pushing, and also realizing the edge system deployment and standardization.

[0175] At the same time, most of the current neural network-based image filtering models in the related art are highly dependent on third-party inference frameworks, and are difficult to integrate, which often requires the introduction of additional libraries, thereby causing problems such as data type conversion, memory copying and inference delay increase between the model and the encoder, making it difficult for the neural network module to be efficiently integrated in a delay-sensitive real-time video coding scene, and even possibly hindering system-level integration due to platform compatibility problems. The technical solution provided in the embodiments of the present application can improve real-time inference efficiency by directly exporting the trained model parameters and embedding them into a C++ encoder, optimizing the convolution calculation process through a block-level parallel computing strategy and an instruction parallel execution strategy, thereby reducing the dependence on third-party deep learning inference frameworks.

[0176] In summary, the technical solution provided in this application demonstrates significant advantages in network architecture design, training strategy, and information utilization strategy. The first model provided in this application, in addition to possessing a lightweight neural network structure and maintaining extremely low computational complexity, can fully exploit the mapping relationship between the reconstructed image and the original image, thereby improving filtering accuracy and image reconstruction quality. Furthermore, the use of a knowledge distillation mechanism to guide the student model to learn the image filtering knowledge of the teacher model enhances the generalization ability of the first model and improves the stability of training the student model. Thus, the image processing method provided in this application shows improvements in complexity control, coding performance, and practical deployment adaptability.

[0177] The technical solution provided in this application can address the contradiction between limited bandwidth and high-quality video services faced by video streaming platforms. It can reduce compression artifacts, improve video quality, and increase encoding efficiency during video encoding and decoding, thereby enabling video platforms to transmit higher resolution and smoother videos under the same bandwidth conditions, thus improving the user's video viewing experience. On the other hand, from an operational cost perspective, the lighter first model can reduce the consumption of server data processing resources, thereby enabling real-time loop filtering processing. This provides real-time video data encoding and decoding support for fields such as 5G+ ultra-high-definition video, cloud VR and augmented reality (AR) metaverse, as well as sports live streaming and automatic editing.

[0178] This application also provides a video processing system. Figure 7 This is a schematic diagram of the structure of the video processing system provided in the embodiments of this application, such as... Figure 7 As shown, the video processing system 700 includes:

[0179] Input unit 701 is used to acquire video;

[0180] The processing unit 702 is used to process the video to obtain a first image;

[0181] A loop filter 703 is used to filter a first image to obtain a second image. The loop filter includes a first model, which includes a first convolutional unit, a nonlinear processing unit, and a second convolutional unit.

[0182] The encoding unit 704 is used to encode the second image to obtain a data stream.

[0183] For example, the input unit can be Figure 6The video input unit in the video input unit; the processing unit can include a frame prediction module, a first residual unit, a transform / quantization unit, an inverse transform / inverse quantization unit and an entropy coding unit, and a second residual unit; the encoding unit can include an entropy coding unit.

[0184] In some embodiments, the first model is distilled by the second model, and the first model is composed of a first convolution unit, a nonlinear processing unit and a second convolution unit arranged in cascade.

[0185] In some embodiments, the loop filter 703 is configured to cache the second image; wherein the second image is a reference image for inter prediction.

[0186] In some embodiments, the first convolution unit is configured to extract first features; wherein the first features include edge and / or texture information of the first image.

[0187] The nonlinear processing unit is configured to process the first features to obtain second features.

[0188] The second convolution unit is configured to fuse and / or reconstruct the second features to obtain the second image with enhanced brightness.

[0189] In some embodiments, the first model is trained by a first loss and a second loss, wherein the first loss is a pixel reconstruction loss, and the second loss is used to represent a distillation loss.

[0190] In some embodiments, the first feature map includes a feature map output by a first layer of the first model, and the second feature map includes a feature map output by a first layer of the second model.

[0191] And / or,

[0192] The first feature map includes a feature map output by a last layer of the first model, and the second feature map includes a feature map output by a last layer of the second model.

[0193] Embodiments of the present application also provide an electronic device, Figure 8 The structure schematic diagram of the electronic device provided by the embodiments of the present application is shown in Figure 8 As shown in the figure, the electronic device 800 includes a processor 801 and a memory 802; the memory 802 stores a computer program; when the computer program is executed by the processor 801, the image processing method as any one of the preceding embodiments can be realized.

[0194] The embodiment of the present application further provides a computer readable storage medium, the storage medium comprising a computer program; the computer program is executed by the processor, and the image processing method in any one of the preceding embodiments can be realized.

[0195] The embodiment of the present application further provides a computer program product, the program product comprising a computer program; the computer program is executed by the processor, and the image processing method in any one of the preceding embodiments can be realized.

[0196] The above description of the various embodiments tends to emphasize the differences between the various embodiments, and the same or similar parts can be referred to each other, and for the sake of brevity, the description is not repeated herein.

[0197] The methods disclosed in the various method embodiments of the present application can be combined arbitrarily without conflict to obtain new method embodiments.

[0198] The features disclosed in the various product embodiments of the present application can be combined arbitrarily without conflict to obtain new product embodiments.

[0199] The features disclosed in the various method or device embodiments of the present application can be combined arbitrarily without conflict to obtain new method or device embodiments.

[0200] It should be noted that the computer readable storage medium above can be a Read Only Memory (ROM), a Programmable Read-Only Memory (PROM), an Erasable Programmable Read-Only Memory (EPROM), an Electrically Erasable Programmable Read-Only Memory (EEPROM), a Ferromagnetic Random Access Memory (FRAM), a Flash Memory, a magnetic surface memory, an optical disc, or a Compact Disc Read-Only Memory (CD-ROM) memory, etc. It can also be various electronic devices including one or any combination of the above memories, such as a mobile phone, a computer, a tablet device, a personal digital assistant, etc. It should be noted that in this paper, the term "include", "contain" or any other variant thereof is intended to cover non-exclusive inclusion, so that the process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or includes elements inherent to such process, method, article or device. Without more limitations, the element defined by the statement "including a" does not exclude the presence of other identical elements in the process, method, article or device including the element. The above application embodiment serial number is only for description, not representing the pros and cons of the embodiments.

[0201] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-mentioned embodiment method can be realized by software and necessary general hardware nodes, of course, it can also be realized by hardware, but in many cases the former is a better implementation. Based on such understanding, the technical solutions of the present application can be embodied in the form of a software product in essence or in the form of a software product that contributes to the prior art. The computer software product is stored in a storage medium (such as ROM / RAM, magnetic disc, optical disc), and includes a plurality of instructions for making a terminal device (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) execute the method described in each embodiment of the present application.

[0202] The computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart block or blocks. Figure 1 one or more flowcharts and / or blocks. Figure 1 one or more flowcharts and / or blocks.

[0203] These computer program instructions can also be stored in a computer readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer readable memory produce an article of manufacture including instructions which implement the function specified in the flowchart block or blocks. Figure 1 one or more flowcharts and / or blocks. Figure 1 one or more flowcharts and / or blocks.

[0204] The computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart block or blocks. Figure 1 one or more flowcharts and / or blocks. Figure 1 one or more flowcharts and / or blocks.

[0205] The above merely provides the preferred embodiment of the present application, and does not limit the patent scope of the present application, and any equivalent structure or equivalent flow transformation made by using the content of the present application specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present application.

Claims

1. An image processing method, characterized in that, include: Get the first image; The first image is filtered using a loop filter to obtain the second image; The loop filter includes a first model, which includes a first convolutional unit, a nonlinear processing unit, and a second convolutional unit.

2. The method according to claim 1, characterized in that, The first model is obtained by distilling the second model, and the first model consists of the first convolutional unit, the nonlinear processing unit and the second convolutional unit arranged in a cascaded manner.

3. The method according to claim 1 or 2, characterized in that, Before acquiring the first image, the method further includes: The video is processed to obtain the first image; After obtaining the second image, the method further includes: The second image is cached; wherein the second image is a reference image for inter-frame prediction.

4. The method according to any one of claims 1 to 3, characterized in that, in: The first convolutional unit is used to extract a first feature; wherein, the first feature includes edge and / or texture information of the first image; The nonlinear processing unit is used to process the first feature to obtain the second feature; The second convolutional unit is used to fuse and / or reconstruct the second feature to obtain the second image with enhanced brightness.

5. The method according to any one of claims 1 to 4, characterized in that, The first model is trained using a first loss and a second loss, wherein the first loss is a pixel reconstruction loss and the second loss is used to represent a distillation loss.

6. The method according to claim 5, characterized in that, The second loss is obtained based on the first feature map and the second feature map, wherein: The first feature map includes the feature map output from the first layer of the first model, and the second feature map includes the feature map output from the first layer of the second model. And / or, The first feature map includes the feature map of the last layer output of the first model, and the second feature map includes the feature map of the last layer output of the second model.

7. A video processing system, characterized in that, include: Input unit, used to acquire video; A processing unit is used to process the video to obtain a first image; A loop filter is used to filter the first image to obtain a second image. The loop filter includes a first model, which includes a first convolutional unit, a nonlinear processing unit, and a second convolutional unit. An encoding unit is used to encode the second image to obtain a data stream.

8. An electronic device, characterized in that, The electronic device includes a processor and a memory; the memory stores a computer program; when the computer program is executed by the processor, it can implement the image processing method as described in any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that, The storage medium includes a computer program; when the computer program is executed by the processor, it is capable of implementing the image processing method as described in any one of claims 1 to 6.

10. A computer program product, characterized in that, The program product includes a computer program; when the computer program is executed by the processor, it is capable of implementing the image processing method as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Image processing method and device, electronic equipment and readable storage medium

    CN113674159A

  • Video coding apparatus, video coding method, video coding program, and non-transitory recording medium

    US20240187579A1