Filtering method of video frame, and coding method, codec and storage medium

By selecting appropriate filtering processing methods according to the frame type of the video frame, including neural network filtering, the problem of poor filtering effect in the existing technology is solved, the image quality of the video frame is improved and the complexity of the neural network is reduced.

CN114157869BActive Publication Date: 2025-10-17ZHEJIANG DAHUA TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202111162686.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-09-30
Publication Date
2025-10-17
Estimated Expiration
2041-09-30

AI Technical Summary

Technical Problem

The existing video frame filtering method is single, resulting in poor filtering processing effect and affecting the image quality of the video frame.

Method used

Different filtering processing methods are determined according to the frame type of the video frame, including neural network filtering, deblocking filtering, sample adaptive compensation filtering and adaptive loop filtering. By determining the filtering processing method corresponding to the frame type of each target image frame, the pixel points of the target image frame are filtered using the filtering processing method that better matches the frame type.

Benefits of technology

The pertinence of filtering is improved, the image quality of the target image frame is improved, and the complexity of neural network filtering is reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114157869B_ABST
    Figure CN114157869B_ABST
Patent Text Reader

Abstract

The application discloses a filtering method of a video frame, and a corresponding coding and decoding method, device and storage medium. The filtering method comprises the following steps: acquiring a target image frame; determining a filtering processing mode corresponding to a frame type of the target image frame; and filtering pixel values of pixel points of the target image frame by using the corresponding filtering processing mode. The filtering processing mode comprises a plurality of filtering sub-steps performed in sequence. The plurality of filtering sub-steps comprise neural network filtering. The filtering processing modes corresponding to at least two different frame types are different. The filtering effect can be improved by using the method.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of video compression coding, in particular to a video frame filtering method, a coding and decoding method, a coder and decoder and a storage medium. BACKGROUND

[0002] A video is a sequence of continuous images, which is composed of continuous video frames. Due to the high similarity between the continuous video frames, in order to facilitate storage and transmission, we need to encode and compress the original video to remove the spatial and temporal redundancy, so as to reduce the network bandwidth in the video transmission process and reduce the storage space occupation. However, the encoding and compression of the video frame will have a certain impact on the quality of the video frame, and often the image frame based on the residual reconstruction has a lower quality than the video frame before encoding and compression.

[0003] At present, filtering the reconstructed image frame is one of the means to improve the quality. The filtering methods include, for example, using a neural network to perform neural network filtering, deblocking filtering, sample adaptive offset filtering and adaptive loop filtering on the image frame. However, the existing filtering methods are relatively single, which also makes the filtering effect not good. SUMMARY

[0004] The present application provides a video frame filtering method, a coding and decoding method, a coder and decoder and a storage medium.

[0005] The first aspect of the present application provides a video frame filtering method, which comprises: obtaining a target image frame; determining a filtering processing mode corresponding to the frame type of the target image frame; and filtering the pixel value of the pixel point of the target image frame using the corresponding filtering processing mode, wherein the filtering processing mode comprises a plurality of filtering sub-steps performed in sequence, the plurality of filtering sub-steps include neural network filtering, and the filtering processing modes corresponding to at least two different frame types are different.

[0006] Therefore, by determining the filtering processing mode corresponding to the frame type of each target image frame, and by limiting the filtering processing modes corresponding to at least two different frame types to be different, the filtering processing mode that is more matched with the frame type can be used to filter the pixel value of the pixel point of the target image frame, which improves the pertinence of filtering, helps to improve the filtering effect and improves the quality of the target image frame. In addition, by determining the filtering processing mode corresponding to the frame type of the target image frame, the complexity of the neural network for performing neural network filtering on each target image frame can be reduced.

[0007] The second aspect of the present application provides an encoding method, which comprises: obtaining a target image frame of a current frame by using a residual corresponding to the current frame; performing the filtering method of the video frame described in the first aspect on the target image frame to obtain a filtered target image frame; and encoding the current frame based on the filtered target image frame.

[0008] The third aspect of the present application provides a decoding method, which comprises: obtaining a target image frame of a current frame by using a residual corresponding to the current frame; performing the filtering method of the video frame described in the first aspect on the target image frame to obtain a filtered target image frame.

[0009] The fourth aspect of the present application provides an encoder, which comprises a processor and a memory coupled to each other, wherein the processor is configured to execute the encoding method described in the second aspect.

[0010] The fifth aspect of the present application provides a decoder, which comprises a processor and a memory coupled to each other, wherein the processor is configured to execute the decoding method described in the third aspect.

[0011] The sixth aspect of the present application provides a computer readable storage medium, which stores a computer program capable of being executed by a processor, and the computer program is configured to implement the filtering method of the video frame described in the first aspect, or the encoding method described in the second aspect, or the decoding method described in the third aspect.

[0012] The above scheme can use the filtering processing mode more matched with the frame type to filter the pixel value of the pixel point of the target image frame, improve the pertinence of filtering, help to improve the filtering effect, and improve the picture quality of the target image frame. BRIEF DESCRIPTION OF DRAWINGS

[0013] Figure 1 is a flowchart of a first embodiment of the filtering method of the video frame of the present application;

[0014] Figure 2 is another flowchart of an embodiment of the filtering method of the video frame of the present application;

[0015] Figure 3 is a filtering processing mode diagram of an embodiment of the filtering method of the video frame of the present application;

[0016] Figure 4 is a flowchart of a second embodiment of the filtering method of the video frame of the present application;

[0017] Figure 5is a schematic diagram of a target image frame and its corresponding reference frame as input in an embodiment of the video frame filtering method of the present application;

[0018] Figure 6 is a flowchart of a third embodiment of the video frame filtering method of the present application;

[0019] Figure 7 is a schematic diagram of filtering the luminance component and two chrominance components of a pixel point of a target image frame in an embodiment of the video frame filtering method of the present application;

[0020] Figure 8 is a flowchart of a fourth embodiment of the video frame filtering method of the present application;

[0021] Figure 9 is a flowchart of a fifth embodiment of the video frame filtering method of the present application;

[0022] Figure 10 is a flowchart of an embodiment of the video frame encoding method of the present application;

[0023] Figure 11 is a flowchart of an embodiment of the video frame decoding method of the present application;

[0024] Figure 12 is a framework diagram of an embodiment of the encoder of the present application;

[0025] Figure 13 is a framework diagram of an embodiment of the decoder of the present application;

[0026] Figure 14 is a framework diagram of an embodiment of the computer readable storage medium of the present application. DETAILED DESCRIPTION

[0027] The schemes of the embodiments of the present application will be described in detail below with reference to the accompanying drawings.

[0028] In the following description, specific details such as specific system structures, interfaces, techniques, etc. are presented in order to provide a thorough understanding of the present application. However, it will be apparent to those skilled in the art that the present application can be practiced without these specific details.

[0029] The terms “system” and “network” are often used interchangeably herein. The term “and / or” herein merely describes an associated relationship, which means that there can be three relationships, for example, A and / or B, which means that A exists alone, A and B exist together, and B exists alone. In addition, the character “ / ” herein generally represents an “or” relationship between the front and rear associated objects. In addition, “multiple” herein means two or more than two.

[0030] In the present application, the meanings of video frame and image frame are the same.

[0031] Please refer to Figure 1 , Figure 1 is a flowchart of a first embodiment of the video frame filtering method of the present application. Specifically, it can include the following steps:

[0032] Step S11: Obtain a target image frame.

[0033] After prediction, the current frame will obtain a corresponding prediction value and a residual corresponding to the current frame. After transformation and quantization, the residual can be combined with the prediction value to obtain an image frame corresponding to the current frame, which is the target image frame in the present application. Specifically, the target image frame can be obtained at the encoding side or at the decoding side, that is, the video frame filtering method of the present application can be applied to the encoding side and the decoding side.

[0034] Step S12: Determine the filtering processing mode corresponding to the frame type of the target image frame.

[0035] The frame type of the target image frame is the same as the frame type of the corresponding current frame. The frame type includes, for example, an intra-coded picture (hereinafter referred to as I frame), a predictive-coded picture (hereinafter referred to as P frame), and a bidirectionally predicted picture (hereinafter referred to as B frame).

[0036] In one specific embodiment, the frame type of the B frame can further include a further classification according to the temporal layer of the B frame, for example, the B frame includes a B1 frame in a first temporal layer, a B2 frame in a second temporal layer, and a B3 frame in a third temporal layer.

[0037] After determining the frame type of the target image frame, the filtering processing mode corresponding to the frame type of the target image frame can be determined. For example, if the frame type of the target image frame is a P frame, the corresponding filtering processing mode is the filtering processing mode corresponding to the P frame. For another example, if the frame type of the target image frame is a B frame, the corresponding filtering processing mode is the filtering processing mode corresponding to the B frame.

[0038] Step S13: Filter the pixel value of the pixel point of the target image frame using the corresponding filtering processing mode.

[0039] In the present application, the filtering processing manner includes a plurality of filtering sub-steps performed in sequence, and the plurality of filtering sub-steps include neural network filtering. The filtering sub-steps include, for example, deblocking filtering, sample adaptive offset filtering, adaptive loop filtering, and neural network filtering. Moreover, the filtering processing manners corresponding to at least two different frame types are different. That is, in the present application, there are at least two different filtering processing manners. For example, the filtering processing manner corresponding to each frame type is different from the filtering processing manners corresponding to other frame types. For another example, the filtering processing manners corresponding to some frame types are different from the filtering processing manners corresponding to other frame types. By filtering the pixel values of the pixel points of the target image frame using the corresponding filtering processing manner, the filtering of the video frame can be implemented.

[0040] By filtering the pixel values of the pixel points of the target image frame, the pixel values of the pixel points can be changed. For example, the pixel value of a certain pixel point of the target image frame is A value under YUV coding, and after filtering, it can be B value. Specifically, the luminance component value and / or the color component value of the A value can be modified to obtain the B value.

[0041] In an embodiment, the neural network filtering is performed by using a neural network containing a skip connection. The neural network containing a skip connection is, for example, a residual network. By performing the neural network filtering by using the neural network containing a skip connection, the filtering effect can be improved by using the better feature learning performance of the neural network containing a skip connection.

[0042] Therefore, by determining the filtering processing manner corresponding to the frame type of each target image frame, and by limiting the filtering processing manners corresponding to at least two different frame types to be different, the filtering processing manner that is more matched with the frame type can be used to filter the pixel values of the pixel points of the target image frame, the targeting of the filtering is improved, which helps to improve the filtering effect and improve the quality of the target image frame. In addition, by determining the filtering processing manner corresponding to the frame type of the target image frame, the complexity of the neural network for performing the neural network filtering on each target image frame can be reduced.

[0043] In an embodiment, among the above-mentioned plurality of filtering sub-steps, the filtering sub-step before the filtered image frame is used as the reference frame can be loop filtering. For example, the plurality of filtering sub-steps include deblocking filtering, sample adaptive offset filtering, adaptive loop filtering, and neural network filtering performed in sequence, and the image frame obtained by the adaptive loop filtering is used as the reference frame. At this time, the deblocking filtering, the sample adaptive offset filtering, and the adaptive loop filtering can be loop filtering.

[0044] In one specific implementation, in a case where the frame type of the target image frame is determined as an I frame, the loop filtering is set to include the neural network filtering. In a case where the frame type of the target image frame is determined as a B frame, the loop filtering is set to not include the neural network filtering. At this time, the neural network filtering is filtering on the B frame as the reference frame, and at this time, the neural network filtering can be considered as post-processing.

[0045] In one specific implementation, in a case where the frame type of the target image frame is determined as an I frame, the loop filtering is set to not include the neural network filtering, and at this time, the neural network filtering is filtering on the I frame as the reference frame, and at this time, the neural network filtering can be considered as post-processing. In a case where the frame type of the target image frame is determined as a B frame, the loop filtering includes the neural network filtering.

[0046] In one specific implementation, in a case where the loop filtering includes the neural network filtering, the loop filtering specifically includes the neural network filtering, and at most two of the deblocking filtering, the sample adaptive offset filtering, and the adaptive loop filtering. For example, the loop filtering includes the neural network filtering, the deblocking filtering, and the sample adaptive offset filtering. For another example, the loop filtering includes the neural network filtering and the deblocking filtering.

[0047] Therefore, by setting the specific filtering processing for the image frames of different frame types, the filtering is improved in pertinence.

[0048] In one implementation, in a case where the frame type of the target image frame is determined as an I frame, the filtering processing mode corresponding to the I frame mentioned in the above step includes: sequentially performing the sample adaptive offset filtering, the neural network filtering, and the adaptive loop filtering on the pixel value of the pixel point of the target image frame. In one specific implementation, the image obtained after the adaptive loop filtering can be used as a reference frame image for subsequent video frame encoding, so as to improve the image quality of the reference frame, and help to improve the image quality of the video frame referring to the reference frame.

[0049] In one implementation, in a case where the frame type of the target image frame is determined as a B frame, the filtering processing mode corresponding to the B frame mentioned in the above step includes: sequentially performing the deblocking filtering, the sample adaptive offset filtering, the adaptive loop filtering, and the neural network filtering on the pixel value of the pixel point of the target image frame. In one specific implementation, the image obtained after the adaptive loop filtering can be used as a reference frame image, and the image obtained after the neural network filtering can be used as an output image for output.

[0050] Therefore, by setting the specific filtering processing mode for the I frame and the B frame, the filtering on the two types of video frames is more targeted, which helps to improve the filtering effect and improve the image quality of the target image frame.

[0051] Referring to Figure 2 , Figure 2 is another flowchart of an embodiment of the video frame filtering method of the present application. In the embodiment, the "determining a filtering processing mode corresponding to the frame type of the target image frame" mentioned in the above steps can include steps S121 and S122.

[0052] Step S121: In the case where it is determined that the frame type of the target image frame is a B frame, determining the temporal layer of the B frame target image frame.

[0053] Because the B frame target image frame can be further subdivided according to the temporal layer it is in, in the case where it is determined that the frame type of the target image frame is a B frame, the temporal layer of the B frame target image frame can be further determined.

[0054] Step S122: Determining the filtering processing corresponding to the temporal layer of the B frame target image frame.

[0055] After the temporal layer of the B frame target image frame is determined, the filtering processing corresponding to the temporal layer of the B frame target image frame can be determined accordingly, so as to realize the targeted filtering of B frames in different temporal layers and improve the filtering effect.

[0056] In an embodiment, in the case where it is determined that the target image frame is a B frame of a first preset temporal layer, the filtering processing corresponding to the temporal layer of the B frame target image frame mentioned above includes deblocking filtering, sample adaptive offset filtering, adaptive loop filtering, and neural network filtering, and the image frame output by the neural network filtering is taken as the reference frame. The first preset temporal layer is, for example, a B1 target image frame in the first temporal layer. In other embodiments, the first preset temporal layer can also be a B frame target image frame in other temporal layers. In addition, taking the image frame output by the neural network filtering as the reference frame means that after the filtering processing corresponding to the B frame of the first preset temporal layer is completed, the image frame obtained by filtering is taken as the reference frame before the processing is completed.

[0057] In an embodiment, in the case where it is determined that the target image frame is a B frame of a second preset temporal layer, the filtering processing corresponding to the temporal layer of the B frame target image frame mentioned above includes deblocking filtering, neural network filtering, sample adaptive offset filtering, and adaptive loop filtering, and the image frame output by the adaptive loop filtering is taken as the reference frame. The second preset temporal layer is, for example, a B2 target image frame in the second temporal layer. In other embodiments, the second preset temporal layer can also be a B frame target image frame in other temporal layers. In addition, taking the image frame output by the adaptive loop filtering as the reference frame means that after the filtering processing corresponding to the B frame of the second preset temporal layer is completed, the image frame obtained by filtering is taken as the reference frame before the processing is completed.

[0058] In one embodiment, when the target image frame is determined to be a B frame of a third preset temporal layer, the filter processing corresponding to the temporal layer of the B frame target image frame includes deblocking filtering, sample adaptive offset filtering, adaptive loop filtering, and neural network filtering. The image frame output by the adaptive loop filtering is used as a reference frame. The third preset temporal layer is, for example, a B3 target image frame in a third temporal layer. In other embodiments, the third preset temporal layer can also be a B frame target image frame in another temporal layer.

[0059] In this embodiment, when the image frame filtered is used as a reference frame, it can also be considered as a step included in the filter processing method.

[0060] Therefore, by determining specific filter processing methods for B frames in different temporal layers, targeted filtering of B frame target image frames in different temporal layers can be achieved to improve the filter effect.

[0061] Referring to Figure 3 , Figure 3 is a schematic diagram of a filter processing method for a video frame in an embodiment of the present application. In Figure 3 , I is an I frame, and B is a B frame, where B1 is a B frame in a first temporal layer, B2 is a B frame in a second temporal layer, and B3 is a B frame in a third temporal layer. The video frames in the solid line boxes are target image frames, and the video frames in the dashed line boxes are video frames filtered and used as reference frames. Specifically, in Figure 3 , SAO is sample adaptive offset, NN is neural network filtering, ALF is adaptive loop filtering, and DBF is deblocking filtering. In addition, in FIG. 3, the dashed lines represent the reference relationship between frames, and the dashed arrows point to the reference frames.

[0062] In one embodiment, the step mentioned above of “filtering the pixel values of the pixel points of the target image frame using the corresponding filter processing method” specifically includes performing neural network filtering on the target image frame and the reference frame of the target image frame as inputs to filter the pixel values of the pixel points of the target image frame using the reference frame. That is, when performing neural network filtering, the target image frame and the reference frame of the target image frame are input to the neural network to filter the pixel values of the pixel points of the target image frame using the reference frame. That is, the target image frame and the reference frame of the target image frame are both input to the neural network to extract feature information of the target image frame and the reference frame of the target image frame by the neural network, and the filtered target image frame is output by the neural network, thereby achieving neural network filtering of the pixel values of the pixel points of the target image frame using the reference frame.

[0063] Specifically, the target image frame and the reference frame of the target image frame can be taken as input, the pixel values of the target image frame and the reference frame of the target image frame can be taken as feature channel information, and then the feature channel information can be taken as input and input into the neural network for neural network filtering.

[0064] In one specific embodiment, the P-frame target image frame and the reference frame thereof can be taken as input of the neural network, the neural network can then extract feature information of the P-frame target image frame and the reference frame thereof, and finally output the filtered P-frame target image frame. In another specific embodiment, the B-frame target image frame and the reference frame thereof can be taken as input of the neural network, the neural network can then extract feature information of the B-frame target image frame and the reference frame thereof, and finally output the filtered B-frame target image frame.

[0065] Therefore, by taking the target image frame and the reference frame thereof as input, the temporal information contained in the reference frame can be utilized when performing neural network filtering, which helps to improve the filtering effect.

[0066] Referring to Figure 4 , Figure 4 is a flowchart of a second embodiment of the video frame filtering method. In this embodiment, the step of "taking the target image frame and the reference frame thereof as input for neural network filtering" specifically includes step S21 and step S22.

[0067] Step S21: determining the number of I-frames in the reference frame of the B-frame target image frame of which the frame type is B-frame.

[0068] For the B-frame target image frame, two reference frames are contained. Therefore, the number of I-frames in the reference frame of the B-frame target image frame can be determined first, so as to perform targeted neural network filtering. Specifically, the number of I-frames in the reference frame of the B-frame target image frame can be two frames, one frame, or no I-frame reference frame.

[0069] Step S22: taking the B-frame target image frame containing different numbers of I-frame reference frames and the corresponding reference frames thereof as input, and filtering the B-frame target image frame using a neural network filtering processing mode corresponding to the number of I-frames in the reference frame of the B-frame target image frame.

[0070] In the embodiment, the neural network filtering processing can be performed according to the number of I frames in the reference frame of the target image frame of the B frame. Specifically, the neural network filtering can be performed on the target image frame of the B frame with different number of I frames in the reference frame by using different neural networks. In the present application, the different neural networks can be different in structure, or the same in structure but different in network parameters, or different in both structure and network parameters.

[0071] Specifically, a neural network filtering processing mode can be determined for the target image frame of the B frame with one I frame in the reference frame, another neural network filtering processing mode can be determined for the target image frame of the B frame with two I frames in the reference frame, and a neural network filtering processing mode can be determined for the target image frame of the B frame with no I frame in the reference frame. In this way, the target image frame of the B frame with different number of I frames in the reference frame can be filtered.

[0072] In one specific embodiment, when the reference frame includes an I frame, the pixel value of the I frame can be used as the channel value of a first preset channel for neural network filtering processing. The specific number of channels of the first preset channel can correspond to the number of pixel value types of the I frame. For example, if the pixel value types of the I frame include a luminance component and two chrominance components, the specific number of channels of the first preset channel is three. For another example, if the pixel value types of the I frame include a luminance component, the specific number of channels of the first preset channel is one. The first preset channel can be a specific channel of a preset channel sequence number. For example, when the first preset channel specifically includes three channels, the first preset channel can include a first channel, a second channel and a third channel, or the last three channels, and the like. The present application does not specifically limit the first preset channel. In one example, after the pixel value of one of the I frames is used as the channel value of the first channel, the pixel value of the target image frame can be used as the channel value of the subsequent channel, and the pixel value of the other I frame can be used as the channel value of the remaining channel. For example, the reference frame of the target image frame of the B frame is one I frame and one B frame. The pixel value of each frame includes a luminance component value and two chrominance component values. The luminance component value and the two chrominance component values of the I frame reference frame can be used as the first channel value to the third channel value, the luminance component value and the two chrominance component values of the target image frame of the B frame can be used as the fourth channel value to the sixth channel value, and the luminance component value and the two chrominance component values of the B frame reference frame can be used as the seventh channel value to the ninth channel value.

[0073] For another example, the pixel value of the I frame reference frame includes a luminance component value and two chrominance component values. The luminance component value and the two chrominance component values of the I frame reference frame can be used as the first channel value, the second channel value and the third channel value, respectively. If only the luminance component value of the reference frame is needed, the luminance component value of the reference frame can be used as the first channel value.

[0074] The reference frame includes an I frame, and the number of I frames can be 1 or 2. In the case of 1 I frame, the pixel value of the I frame reference frame can be used as the channel value of the first preset channel of the channel information input into the neural network. In the case of 2 I frames, the pixel value of the I frame reference frame in the time sequence in front can be used as the channel value of the first preset channel of the channel information input into the neural network, and the pixel value of the other I frame reference frame can be used as the last channel value.

[0075] In one embodiment, when it is determined that the reference frames are all B frames, the pixel value of the B frame reference frame in the lowest time domain layer can be used as the channel value of the second preset channel for neural network filtering processing. The specific number of channels of the second preset channel can correspond to the number of pixel value types of the B frame. For the specific setting method of the second preset channel, please refer to the related description of the first preset channel above, which will not be repeated here. If the time domain layers of the B frame reference frames are the same, the B frame in the time sequence in front can be further used as the channel value of the second preset channel for neural network filtering processing. The present application does not make specific limitations on the second preset channel.

[0076] In one embodiment, when it is determined that the reference frames are B frames and P frames, the pixel value of the P frame reference frame can be used as the first channel value for neural network filtering processing.

[0077] In one embodiment, the target image frame of the P frame type can use the pixel value of its reference frame as the first channel value.

[0078] Therefore, when the target image frame and the reference frame of the target image frame are used as input for neural network filtering, the specific determination method of the channel information input into the neural network helps to improve the filtering effect of the neural network.

[0079] Reference is made to Figure 5 , Figure 5 FIG. 1 is a schematic diagram of the present application in which the target image frame and its corresponding reference frame are used as input in the embodiment of the video frame filtering method of the present application. In Figure 5 , the video frames in the solid line box are target image frames, in which I is an I frame, and B is a B frame, in which B1 is a B frame in the first time domain layer, B2 is a B frame in the second time domain layer, and B3 is a B frame in the third time domain layer. The numbers in the brackets on each image frame represent the encoding order of the frame. The video frames in the dashed line box are the video frames after filtering and used as reference frames. NN1, NN2 and NN3 are three different neural network filtering processes. In Figure 3For the B3(5) target image frame, the corresponding neural network filtering processing is NN2, the input of the NN2 includes the I(1) target image frame, the B3(5) target image frame, and the B2(4) target image frame, wherein the I(1) target image frame and the B2(4) target image frame are reference frames of the B3(5) target image frame.

[0080] In an embodiment, the "filtering the pixel value of the pixel point of the target image frame using the corresponding filtering processing mode" mentioned in the above steps can specifically include: filtering the luminance component and the chrominance component of the pixel point of the target image frame in the same neural network filtering processing; or filtering the luminance component and the chrominance component of the pixel point of the target image frame in different neural network filtering processes.

[0081] In the embodiment, filtering the luminance component and the chrominance component of the pixel point of the target image frame in the same neural network filtering processing can be considered as filtering the luminance component and the chrominance component of the pixel point of the target image frame by using one neural network. If the reference frame is also input to the neural network, it can be considered as filtering the luminance component and the chrominance component of the pixel point of the target image frame and the reference frame by using one neural network. Therefore, by filtering the luminance component and the chrominance component of the pixel point of the target image frame in the same neural network filtering processing, the overall filtering of the target image frame can be realized.

[0082] In the embodiment, filtering the luminance component and the chrominance component of the pixel point of the target image frame in different neural network filtering processes can be considered as filtering the luminance component and the chrominance component of the pixel point of the target image frame by using different neural networks respectively. If the reference frame is also input to the neural network, it can be considered as filtering the luminance component and the chrominance component of the pixel point of the target image frame and the reference frame by using different neural networks. Therefore, by filtering the luminance component and the chrominance component of the pixel point of the target image frame in different neural network filtering processes respectively, the targeted filtering of the luminance component and the chrominance component of the target image frame can be realized, which helps to improve the filtering effect.

[0083] In an embodiment, the "filtering the luminance component and the chrominance of the pixel point of the target image frame in different neural network filtering processes" mentioned above specifically includes: filtering the luminance component and two chrominance components of the pixel point of the target image frame in three neural network filtering processes respectively. Therefore, by filtering the luminance component and two chrominance components of the pixel point of the target image frame in three neural network filtering processes respectively, the respective filtering of the luminance component and two chrominance components of the pixel point of the target image frame can be realized, which improves the targeting of the filtering and helps to improve the filtering effect.

[0084] Referring to Figure 6 , Figure 6 is a flowchart of a third embodiment of the video frame filtering method of the present application. In the embodiment, the above-mentioned "filtering the luminance component and the two chrominance components of the pixel points of the target image frame in three different neural network filtering processes respectively" specifically comprises steps S31 to S33.

[0085] Step S31: performing first neural network filtering process on the luminance component of the pixel points of the target image frame to filter, and taking the chrominance components as the input of the first neural network filtering process to filter the luminance component with the chrominance components.

[0086] In the embodiment, when performing the first neural network filtering process on the luminance component of the pixel points of the target image frame to filter, the luminance component and the chrominance components are simultaneously input to the neural network to perform the first neural network filtering process, so that the neural network can filter the luminance component with the characteristic information of the chrominance components, thereby improving the filtering effect. Specifically, both of the two chrominance components can be taken as the input, or one of the two chrominance components can be taken as the input.

[0087] Step S32: performing second neural network filtering process on one of the chrominance components of the pixel points of the target image frame to filter, and taking the luminance component as the input of the second neural network filtering process to filter the one of the chrominance components with the luminance component.

[0088] In the embodiment, when performing the second neural network filtering process on one of the chrominance components of the pixel points of the target image frame to filter, the luminance component is simultaneously input to the neural network to perform the second neural network filtering process, so that the neural network can filter the one of the chrominance components with the characteristic information of the luminance component, thereby improving the filtering effect.

[0089] Step S33: performing third neural network filtering process on the other of the chrominance components of the pixel points of the target image frame to filter, and taking the luminance component as the input of the third neural network filtering process to filter the other of the chrominance components with the luminance component.

[0090] In the embodiment, when performing the third neural network filtering process on the other of the chrominance components of the pixel points of the target image frame to filter, the luminance component is simultaneously input to the neural network to perform the third neural network filtering process, so that the neural network can filter the other of the chrominance components with the characteristic information of the luminance component, thereby improving the filtering effect.

[0091] Referring toFigure 7 , Figure 7 is a schematic diagram of filtering the luminance component and two chrominance components of a pixel point of a target image frame in an embodiment of the video frame filtering method of the present application. In Figure 7 , Y represents the luminance component of the target image frame, and U and V represent the two chrominance components of the target image frame. In Figure 7 part (a), after the U and V components are deconvolved, the corresponding image is equal to the image of the luminance component. At this time, by taking Y, U and V components as input, input into the residual network ResNet, the filtered Y component can be obtained, and filtering of the Y component by the U and V components is realized. Figure 7 parts (b) and (c) are not described again.

[0092] Referring to Figure 8 , Figure 8 is a flowchart of the fourth embodiment of the video frame filtering method of the present application. In this embodiment, for each target image frame, the above-mentioned "filtering the pixel value of the pixel point of the target image frame using the corresponding filtering processing mode" includes steps S41 and S42.

[0093] Step S41: Obtain the quantization parameters of a plurality of quantization blocks contained in the current frame corresponding to the target image frame.

[0094] In this embodiment, when the current frame is quantized, the current frame is divided to obtain a plurality of blocks for quantization, which are the quantization blocks. In order to further improve the filtering effect, the quantization parameters of a plurality of quantization blocks contained in the current frame corresponding to the target image frame are obtained, and then targeted filtering is performed according to the different quantization parameters of the quantization blocks.

[0095] Step S42: Perform neural network filtering corresponding to the preset range on the blocks in the target image frame corresponding to the quantization blocks with quantization parameters in different preset ranges.

[0096] The corresponding block of the quantization block in the target image frame can be considered as the block corresponding to the position of the quantization block in the target image frame. The preset range can be set as needed. For example, the quantization parameter range is [0, 51], and the quantization parameter can be divided into 6 preset ranges, [0, 24], [25, 29], [30, 34], [35, 39], and [40, 51]. Then, based on the quantization parameter of each quantization block, the corresponding neural network filtering of the preset range can be performed on the block corresponding to the quantization block in the target image frame. For example, the quantization parameter range of a certain quantization block is 45, and it can be determined that it belongs to the [40, 51] preset range. Then, the block corresponding to the quantization block in the target image frame can be determined, and the neural network filtering corresponding to the [40, 51] preset range can be performed on the block corresponding to the quantization block in the target image frame.

[0097] Specifically, a corresponding neural network filtering processing method can be set for each preset range, or the same corresponding neural network filtering processing method can be set for part of the preset ranges. The specific setting can be performed as needed, which is not limited here.

[0098] Therefore, by obtaining the quantization parameters of the quantization blocks contained in the current frame corresponding to the target image frame, and then performing targeted filtering on the block corresponding to the quantization block in the target image frame according to the different quantization parameters of the quantization block, the filtering effect can be improved.

[0099] In an embodiment, for each quantization block, the above-mentioned "performing neural network filtering on the block corresponding to the quantization block in the target image frame according to the different preset ranges of the quantization parameter" can specifically include: performing neural network filtering on the luminance component and the chrominance component of the pixel points of the block corresponding to the quantization block in the target image frame. That is, the luminance component and the chrominance component of the pixel points of the block corresponding to the quantization block in the target image frame can be respectively filtered by a neural network. For example, the luminance component and the two chrominance components of a certain quantization block can be respectively filtered by a neural network.

[0100] In a specific embodiment, the luminance component of the block corresponding to the quantization block in the target image frame can also be filtered by a corresponding neural network filtering, and the chrominance components of the blocks corresponding to all quantization blocks in the target image frame can be filtered by the same neural network filtering processing method. Or, the two chrominance components of the blocks corresponding to all quantization blocks in the target image frame can be respectively filtered by a neural network filtering processing method.

[0101] In this way, on the basis of performing targeted filtering on the block corresponding to the quantized block in the target image frame, different components of the block corresponding to the quantized block in the target image frame can be further filtered to improve the filtering effect.

[0102] See Figure 9 , Figure 9 1 is a flowchart of the fifth embodiment of the method for filtering video frames of the present application. In this embodiment, the above-mentioned step of "filtering the pixel values ​​of the pixel points of the target image frame using the corresponding filtering processing method" specifically includes step S51 and step S52.

[0103] Step S51: Divide the target image frame into a number of target blocks.

[0104] The target image frame may be segmented using a common method in the art. The target block may be, for example, a coding unit (CU).

[0105] In a specific embodiment, the target image frame can be divided into several target blocks including a stitching area. The stitching area is, for example, a coding unit. The size of the target block is not smaller than the stitching area. For example, the size of the target block is larger than the stitching area. By setting the size of the target block to be larger than the stitching area, when the target block is subsequently subjected to neural network filtering, the pixel value information of the pixel points outside the stitching area can be used to filter the stitching area, which helps to improve the filtering effect. In a specific embodiment, if the stitching area is the edge of the target image frame, padding can be performed around the stitching area to obtain the target block.

[0106] Step S52: using a corresponding filtering processing method to filter the pixel values ​​of the pixel points of a plurality of target blocks of the target image frame to obtain filtered target blocks.

[0107] The corresponding filtering processing mode is the filtering processing mode corresponding to the frame type of the target image frame. In this embodiment, the pixel values ​​of the pixel points of several target blocks of the target image frame are filtered to obtain filtered target blocks, thereby achieving filtering of the entire target image frame.

[0108] In one embodiment, after the step of "filtering the pixel values ​​of the pixel points of several target blocks of the target image frame using a corresponding filtering method to obtain a filtered target block," the video frame filtering method of the present application further includes: obtaining a filtered target image frame using the filtered target block. It will be appreciated that the filtered target block is obtained by filtering the target blocks obtained by segmenting the target image frame. Therefore, the filtered target blocks can be correspondingly spliced ​​to obtain a filtered target image frame.

[0109] In one embodiment, for the case that the target image frame is divided into a plurality of target blocks containing the stitching region, the stitching region in the filtered target block can be used for stitching, so as to obtain the filtered target image frame.

[0110] It can be understood that the filtering method of the video frame mentioned in the above embodiments is only an exemplary description, and different filtering methods of the video frame mentioned in the above embodiments can be combined with each other to obtain a new reconstruction frame method, which is not limited in the present application.

[0111] In one embodiment, the filtering method of the video frame of the present application can further include the following steps 1 to 3 to train the neural network used for neural network filtering.

[0112] Step 1: Obtain a target image frame and determine the filtering processing mode corresponding to the frame type of the target image frame.

[0113] For details of the reconstructed frame image, please refer to the related description of the above embodiments, which will not be repeated here.

[0114] In the present embodiment, the filtering processing mode includes a plurality of filtering sub-steps performed in sequence, the neural network filtering is included in some of the filtering sub-steps, and the filtering processing modes corresponding to at least two different frame types are different. For specific determination of the filtering processing mode corresponding to the frame type, please refer to the related description of the above embodiments, which will not be repeated here.

[0115] Step 2: Filter the pixel values of the pixel points of the target image frame using the corresponding filtering processing mode to obtain a filtered target image frame.

[0116] For details, please refer to the related description of the above embodiments, which will not be repeated here.

[0117] Step 3: Adjust the network parameters of the neural network used for neural network filtering based on the difference between the current frame corresponding to the target image frame and the filtered target image frame.

[0118] In the present embodiment, the current frame corresponding to the target image frame is the original video frame without encoding. Specifically, when training the neural network used for neural network filtering in the filtering processing mode corresponding to the I frame, the neural network used for neural network filtering in the filtering processing mode corresponding to the P frame or B frame can be trained using a training set composed of P frames or B frames, respectively.

[0119] In one embodiment, the neural networks for filtering the target image frames of different frame types can be trained in a reference relationship between different types of frames. For example, referring to the filtering method of the video frames shown in FIG. 3, since the I frame is an intra prediction frame, the neural network for filtering the target image frame of the I frame can be trained first, and the target image frame obtained after ALF filtering is taken as the reference frame. Then, the neural network for filtering the target image frame of the B frame with the time domain layer 1 can be trained second, and the target image frame obtained after NN filtering is taken as the reference frame. Then, the neural network for filtering the target image frame of the B frame with the time domain layer 2 can be trained third, and the target image frame obtained after ALF filtering is taken as the reference frame. Finally, the neural network for filtering the target image frame of the B frame with the time domain layer 3 can be trained fourth. In this way, the neural network for filtering the target image frame of the I frame can make the subsequent target image frame with the I frame as the reference frame have higher picture quality, and reduce the error introduced when the I frame is taken as the reference frame, thereby improving the training effect. Similarly, the neural network for filtering the target image frame of the B frame with the time domain layer 1, and the neural network for filtering the target image frame of the B frame with the time domain layer 2 can make the subsequent target image frame with the B1 frame and the B2 frame as the reference frame have higher picture quality, and reduce the error introduced when the B1 frame and the B2 frame are taken as the reference frame, thereby improving the training effect.

[0120] In one embodiment, the neural network for filtering the target image frame of the P frame can be trained after the training of the neural network for filtering the target image frame of the I frame. In this way, the subsequent target image frame with the P frame as the reference frame can have higher picture quality, and the error introduced when the P frame is taken as the reference frame can be reduced, thereby improving the training effect.

[0121] Referring to Figure 10 , Figure 10 is a flowchart of an embodiment of the method for encoding the video frame. In this embodiment, the method for encoding the video frame specifically includes:

[0122] Step S61: obtaining the target image frame of the current frame by using the residual corresponding to the current frame.

[0123] The current frame can be an original video frame to be encoded, and the residual corresponding to the current frame can be obtained by using a common video frame prediction method in the field. For example, after the current frame is predicted by using an intra-frame prediction method, a prediction value can be obtained, and the residual corresponding to the current frame can be obtained by subtracting the prediction value from an actual value of the current frame. The residual can be changed and quantized to obtain a quantized residual, and the target image frame can be obtained based on the quantized residual and the prediction value of the current frame.

[0124] Step S62: performing the video frame filtering method on the target image frame to obtain a filtered target image frame.

[0125] In the embodiment, the video frame filtering method can be the method mentioned in the embodiment of the video frame filtering method, and the specific reconstruction process can be referred to the description of the embodiment, which is not described herein.

[0126] In an embodiment, the video frame filtering method can be performed on the luminance component and the two chrominance components of the target image frame respectively, and then the filtered components are combined to obtain the filtered target image frame.

[0127] Step S63: encoding the current frame based on the filtered target image frame.

[0128] After the filtered target image frame is obtained, the current frame can be encoded by using a common encoding method in the field. For example, a syntax switch related to the video frame filtering method can be encoded, so that the decoding side can determine whether the video frame filtering method needs to be performed according to the syntax switch related to the video frame filtering method.

[0129] In an embodiment, a first switch of the neural network filtering processing switch can be set for a current sequence to which the current frame belongs. When the first switch of the neural network filtering processing switch corresponding to the current sequence to which the current frame belongs is on, it indicates that the video frame sequence can perform the video frame filtering method mentioned in the embodiment. When the first switch of the neural network filtering processing switch corresponding to the current sequence to which the current frame belongs is off, it indicates that the video frame sequence does not perform the video frame filtering method mentioned in the embodiment. Therefore, in a specific embodiment, the step S62 is performed when the first switch is on. Therefore, by setting the neural network filtering processing switch for the current sequence to which the current frame belongs, the control of whether the video frame filtering method is performed on the current sequence can be realized.

[0130] In an embodiment, the encoding of the current frame based on the filtered target image frame includes steps S631 to S63.

[0131] Step S631: compare the quality of the filtered target image frame and the unfiltered target image frame.

[0132] The method of comparing the quality of the filtered target image frame and the unfiltered target image frame can be a common quality comparison method in the art, for example, calculating the rate distortion of the filtered target image frame and the unfiltered target image frame respectively, and then comparing the quality of the filtered target image frame and the unfiltered target image frame.

[0133] In the case of performing the filtering method of the video frame on the luminance component, the two chrominance components of the target image frame, the quality of each component of the filtered target image frame can be compared with the corresponding component of the unfiltered target image frame.

[0134] If the filtered target image frame is better than the unfiltered target image frame, step S632 can be performed, and if the filtered target image frame is worse than the unfiltered target image frame, step S633 can be performed.

[0135] Step S632: set the neural network filtering processing switch of the current frame to on.

[0136] The filtered target image frame is better than the unfiltered target image frame, which means that the filtering method of the video frame can be used to filter the target image frame to improve the quality, so the neural network filtering processing switch of the current frame can be set to on, so that the decoding side can perform the filtering method of the video frame on the current frame based on the neural network filtering processing switch of the current frame.

[0137] In the case of performing the filtering method of the video frame on the luminance component, the two chrominance components of the target image frame, if there is any one component of the filtered target image frame that is better than the corresponding component of the unfiltered target image frame, the neural network filtering processing switch corresponding to the any one component of the current frame can be set to on. For example, the luminance component of the filtered target image frame is better than the luminance component of the unfiltered target image frame, and the neural network filtering processing switch corresponding to the luminance component of the current frame can be set to on.

[0138] Step S633: set the neural network filtering processing switch of the current frame to off.

[0139] The filtered target image frame is worse than the unfiltered target image frame, which means that it is unnecessary to use the filtering method of the video frame to filter the target image frame. Therefore, the neural network filtering processing switch of the current frame can be set to off, so that the decoding side can not perform the filtering method of the video frame on the current frame based on the neural network filtering processing switch of the current frame.

[0140] In the case that the filtering method of the video frame is performed on the luminance component, the two chrominance components of the target image frame, if any one component of the target image frame after filtering is worse than the corresponding component of the target image frame before filtering, the neural network filtering processing switch corresponding to the any one component of the current frame can be set to off. For example, if the luminance component of the target image frame after filtering is worse than the luminance component of the target image frame before filtering, the neural network filtering processing switch corresponding to the luminance component of the current frame can be set to off.

[0141] Therefore, by comparing the quality of the target image frame after filtering and the target image frame before filtering, it can be determined whether the filtering method of the video frame needs to be performed on the target image frame, thereby realizing the control of whether the filtering method of the video frame is performed on the target image frame.

[0142] In an embodiment, the step mentioned above that the filtering method of the video frame is performed on the target image frame to obtain the target image frame after filtering specifically includes: the filtering method of the video frame is performed on a plurality of target blocks obtained by dividing the target image frame to obtain the target image frame after filtering. In this embodiment, the target block is, for example, a coding unit. In other embodiments, the target block can also be set as needed. After performing the filtering method of the video frame on each target block, the target blocks after filtering are spliced to obtain the target image frame after filtering.

[0143] In this case, the step mentioned above that the current frame is encoded based on the target image frame after filtering includes:

[0144] Step S634: comparing the quality of each target block after filtering and the target block before filtering.

[0145] The specific comparison method can refer to the description of step S631 above, which will not be described here again.

[0146] If the target block after filtering is better than the target block before filtering, step S635 can be performed, and if the target block after filtering is worse than the target block before filtering, step S636 can be performed.

[0147] Step S635: setting the neural network filtering processing switch corresponding to the target block to on.

[0148] The target block after filtering being better than the target block before filtering means that the filtering method of the video frame can be used to filter the target block to improve the quality. Therefore, the neural network filtering processing switch corresponding to the target block can be set to on, so that the decoding side can perform the filtering method of the video frame on the target block based on the neural network filtering processing switch of the target block.

[0149] Step S636: set the neural network filtering processing switch corresponding to the target block to off.

[0150] The filtered target block is worse than the target block before filtering, which means that the filtering method of the video frame does not have to be used to filter the target block. Therefore, the neural network filtering processing switch corresponding to the target block can be set to off, so that the decoding side can not perform the filtering method of the video frame on the target block based on the neural network filtering processing switch of the target block.

[0151] Step S637: determine whether there is at least one target block whose corresponding neural network filtering processing switch is set to on.

[0152] For the target image frame, if there is at least one target block whose corresponding neural network filtering processing switch is set to on, it means that the filtering method of the video frame can be performed on the target block whose corresponding neural network filtering processing switch is set to on in the target image frame. Therefore, it can be determined in advance whether there is at least one target block whose corresponding neural network filtering processing switch is set to on.

[0153] If there is at least one target block whose corresponding neural network filtering processing switch is set to on, step S638 can be performed, and if there is no target block whose corresponding neural network filtering processing switch is set to on, step S639 can be performed.

[0154] Step S638: set the neural network filtering processing switch of the current frame to on.

[0155] Because there is at least one target block whose corresponding neural network filtering processing switch is set to on, it means that the filtering method of the video frame can be performed on the target block whose corresponding neural network filtering processing switch is set to on in the target image frame to improve the image quality. Therefore, the neural network filtering processing switch of the current frame can be set to on accordingly, so that the decoding side can perform the filtering method of the video frame on the target block based on the neural network filtering processing switch of the current frame and the neural network filtering processing switch corresponding to the target block.

[0156] Step S639: set the neural network filtering processing switch of the current frame to off.

[0157] Because there is no target block whose corresponding neural network filtering processing switch is set to on, it means that the filtering method of the video frame does not have to be performed on all target blocks in the target image frame, so the neural network filtering processing switch of the current frame can be set to off, so that the decoding side can not perform the filtering method of the video frame on the current frame based on the neural network filtering processing switch of the current frame.

[0158] Therefore, by determining the neural network filtering processing switch corresponding to each target block, whether the video frame filtering method needs to be performed on each target block can be realized, and the control of whether the video frame filtering method is performed on the target block is realized.

[0159] Referring to Figure 11 , Figure 11 is a flowchart of an embodiment of the video frame decoding method of the present application. In the embodiment, the video frame decoding method specifically comprises:

[0160] Step S71: obtaining a target image frame of a current frame by using a residual corresponding to the current frame.

[0161] The residual corresponding to the current frame is obtained after decoding the code stream transmitted from the encoding side, for example. After inverse quantization and inverse transformation are performed on the residual corresponding to the current frame, the target image frame corresponding to the current frame can be obtained in combination with the prediction value of the current frame.

[0162] Step S72: performing a video frame filtering method on the target image frame to obtain a filtered target image frame.

[0163] In the embodiment, the video frame filtering method can be the method mentioned in the above embodiment of the video frame filtering method, and the specific reconstruction process is described in the above embodiment, which will not be described here.

[0164] In one specific embodiment, after decoding the code stream transmitted from the encoding side, the specific switch condition of the neural network filtering processing switch of the current frame can be obtained. In the case that the neural network filtering processing switch of the current frame is on, the video frame filtering method can be performed on the target image frame to obtain the filtered target image frame.

[0165] In one specific embodiment, after decoding the code stream transmitted from the encoding side, the neural network filtering processing switch of the current frame and the neural network filtering processing switches corresponding to a plurality of target blocks obtained after the target image frame corresponding to the current frame is divided can be obtained. In the case that the neural network filtering processing switch of the current frame is on, the video frame filtering method can be performed on the target block whose corresponding neural network filtering processing switch is on. Finally, all the target blocks are spliced to obtain the filtered target image frame.

[0166] Therefore, by using the above decoding method, the video frame filtering method can be performed on the decoding side.

[0167] Referring to Figure 12 , Figure 12is a framework schematic diagram of an embodiment of an encoder of the present application. The encoder 120 comprises a memory 121 and a processor 122 coupled with each other, and the processor 122 is configured to execute program instructions stored in the memory 121 to implement the steps of any of the above-mentioned embodiments of the method for encoding a video frame. In a specific implementation scenario, the encoder 120 can include but is not limited to a microcomputer, a server, and in addition, the encoder 120 can also include a notebook computer, a tablet computer and other mobile devices, which are not limited herein. Specifically, the processor 122 is configured to control itself and the memory 121 to implement the steps of any of the above-mentioned embodiments of the method for encoding a video frame.

[0168] Referring to Figure 13 , Figure 13 is a framework schematic diagram of an embodiment of a decoder of the present application. The decoder 130 comprises a memory 131 and a processor 132 coupled with each other, and the processor 132 is configured to execute program instructions stored in the memory 131 to implement the steps of any of the above-mentioned embodiments of the method for decoding a video frame. In a specific implementation scenario, the decoder 130 can include but is not limited to a microcomputer, a server, and in addition, the decoder 130 can also include a notebook computer, a tablet computer and other mobile devices, which are not limited herein. Specifically, the processor 132 is configured to control itself and the memory 131 to implement the steps of any of the above-mentioned embodiments of the method for decoding a video frame.

[0169] The above-mentioned processor can also be referred to as a CPU (Central Processing Unit). The processor can be an integrated circuit chip with processing capability. The processor can also be a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field-Programmable Gate Array) or other programmable logic device, discrete gate or transistor logic device, discrete hardware component. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor. In addition, the processor can be implemented by an integrated circuit chip.

[0170] Referring to Figure 14 , Figure 14 is a framework schematic diagram of an embodiment of a computer readable storage medium of the present application. The computer readable storage medium 140 stores program instructions 141 capable of being executed by a processor, and the program instructions 141 are configured to implement the steps of any of the above-mentioned embodiments of the method for filtering a video frame, the steps of any of the above-mentioned embodiments of the method for encoding a video frame, or the steps of any of the above-mentioned embodiments of the method for decoding a video frame.

[0171] The above scheme can use the filter processing mode that is more matched with the frame type to filter the pixel value of the pixel point of the target image frame, improve the pertinence of filtering, help to improve the filtering effect, and improve the image quality of the target image frame. In addition, by determining the filter processing mode corresponding to the frame type of the target image frame, the complexity of the neural network for performing neural network filtering on each target image frame can be reduced.

[0172] In some embodiments, the apparatus provided by the embodiments of the present disclosure has functions or includes modules that can be used to perform the methods described in the above method embodiments, and the specific implementation can refer to the description of the above method embodiments. For brevity, details are not repeated here.

[0173] The above description of various embodiments tends to emphasize the differences between various embodiments, and the same or similar parts can be mutually referred to. For brevity, details are not repeated here.

[0174] In several embodiments provided in the present application, it should be understood that the disclosed methods and devices can be implemented in other ways. For example, the above-described device implementation is only schematic, for example, the division of modules or units is only a logical function division, and actual implementation can have another division manner, for example, a unit or component can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the shown or discussed each other can be indirect coupling or communication connection through some interface, device or unit, and can be electrical, mechanical or other forms.

[0175] The unit described as a separate component can or can not be physically separated, and the component shown as a unit can or can not be a physical unit, that is, it can be located in one place, or it can be distributed on a network unit. According to actual needs, part or all of the units can be selected to achieve the purpose of the present embodiment scheme.

[0176] In addition, each functional unit in each embodiment of the present application can be integrated in one processing unit, or each unit can exist physically, or two or more units can be integrated in one unit. The above integrated unit can be realized in the form of hardware or in the form of a software functional unit.

[0177] The integrated unit, if implemented in the form of a software function unit and sold or used as an independent product, can be stored in a computer readable storage medium. Based on such an understanding, the technical solutions of the present application, essentially or in other words, the part that contributes to the prior art or the whole or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, includes a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the methods of the various embodiments of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various media that can store program codes.

Claims

1. A method for filtering a video frame, characterized in that: include: Get the target image frame; Determining a filtering processing method corresponding to a frame type of the target image frame; Filtering pixel values ​​of pixels of the target image frame using the corresponding filtering processing method, wherein the filtering processing method includes a plurality of filtering sub-steps performed in sequence, the plurality of filtering sub-steps including neural network filtering, and at least two different frame types correspond to different filtering processing methods; The filtering of the pixel values ​​of the pixel points of the target image frame using the corresponding filtering processing method includes: Performing neural network filtering on the target image frame and a reference frame of the target image frame as input, so as to perform neural network filtering on pixel values ​​of the target image frame using the reference frame; Among them, the reference frame is an image frame obtained by filtering the several filtering sub-steps performed in sequence; the neural network filtering is performed on the target image frame and the reference frame of the target image frame as input, including: determining the number of I frames in the reference frame of the B-frame target image frame whose frame type is B-frame; taking the B-frame target image frame containing different numbers of I-frame reference frames and its corresponding reference frames as input, and filtering the B-frame target image frame using a neural network filtering processing method corresponding to the number of I frames in the reference frame of the B-frame target image frame.

2. The method according to claim 1, characterized in that The determining of the filtering processing method corresponding to the frame type of the target image frame includes: In the plurality of filtering sub-steps, the filtering sub-step before taking the filtered image frame as the reference frame is performed as loop filtering; When it is determined that the frame type of the target image frame is an I frame, the loop filtering includes a neural network filter; when it is determined that the frame type of the target image frame is a B frame, the loop filtering does not include a neural network filter; or, When it is determined that the frame type of the target image frame is an I frame, the loop filtering does not include neural network filtering; when it is determined that the frame type of the target image frame is a B frame, the loop filtering includes neural network filtering.

3. The method according to claim 2, characterized in that In the case where the loop filtering includes neural network filtering, the loop filtering specifically includes: neural network filtering, and at most two of deblocking filtering, sample adaptive compensation filtering, and adaptive loop filtering.

4. The method according to claim 1, wherein When it is determined that the frame type of the target image frame is an I frame, the corresponding filtering processing method includes: performing sample adaptive compensation filtering, neural network filtering, and adaptive loop filtering on pixel values ​​of pixel points of the target image frame in sequence; When it is determined that the frame type of the target image frame is a B frame, the corresponding filtering processing method includes: performing deblocking filtering, sample adaptive compensation filtering, adaptive loop filtering and neural network filtering on the pixel values ​​of the pixel points of the target image frame in sequence.

5. The method according to claim 1, wherein The determining of the filtering processing method corresponding to the frame type of the target image frame includes: In the case where it is determined that the frame type of the target image frame is a B frame, determining a temporal layer of the B frame target image frame; A filtering process corresponding to the temporal layer of the B-frame target image frame is determined.

6. The method according to claim 5, characterized in that The determining of the filtering process corresponding to the temporal layer of the B-frame target image frame includes: When it is determined that the target image frame is a B frame of the first preset temporal layer, the filtering processing corresponding to the temporal layer of the B frame target image frame includes: deblocking filtering, sample adaptive compensation filtering, adaptive loop filtering and neural network filtering, and the image frame output by the neural network filtering is used as a reference frame; When it is determined that the target image frame is a B frame of the second preset temporal layer, the filtering processing corresponding to the temporal layer of the B frame target image frame includes: deblocking filtering, neural network filtering, sample adaptive compensation filtering, and adaptive loop filtering, and the image frame output by the adaptive loop filtering is used as a reference frame; When it is determined that the target image frame is a B frame of the third preset time domain layer, the filtering processing corresponding to the time domain layer of the B frame target image frame includes: deblocking filtering, sample adaptive compensation filtering, adaptive loop filtering and neural network filtering, and the image frame output by the adaptive loop filtering is used as a reference frame.

7. The method according to any one of claims 1 to 6, characterized in that The filtering of the pixel values ​​of the pixel points of the target image frame using the corresponding filtering processing method includes: dividing the target image frame into a plurality of target blocks; Using the corresponding filtering processing method to filter the pixel values ​​of the pixel points of the plurality of target blocks of the target image frame to obtain filtered target blocks; After filtering the pixel values ​​of the pixel points of several target blocks of the target image frame using the corresponding filtering processing method to obtain filtered target blocks, the method further includes: obtaining a filtered target image frame using the filtered target blocks.

8. The method according to claim 7, characterized in that The dividing the target image frame into a plurality of target blocks includes: dividing the target image frame into a plurality of target blocks including a splicing area; The obtaining of a filtered target image frame by using the filtered target block includes: obtaining a filtered target image frame by using a spliced ​​area in the filtered target block.

9. The method according to claim 1, characterized in that The filtering of the B-frame target image frame using a neural network filtering processing method corresponding to the number of I frames in the reference frame of the B-frame target image frame includes: When it is determined that the reference frame includes an I frame, performing neural network filtering processing using pixel values ​​of one of the I frames as channel values ​​of a first preset channel; When it is determined that all the reference frames are B frames, the pixel values ​​of the B frame reference frame in the lowest temporal layer are used as channel values ​​of the second preset channel to perform neural network filtering processing.

10. The method according to any one of claims 1 to 6, characterized in that The filtering of the pixel values ​​of the pixel points of the target image frame using the corresponding filtering processing method includes: filtering the luminance component and the chrominance component of the pixel points of the target image frame in the same neural network filtering process; or, The luminance component and chrominance component of the pixel points of the target image frame are filtered in different neural network filtering processes.

11. The method according to claim 10, characterized in that The filtering of the brightness component and chrominance of the pixel points of the target image frame in different neural network filtering processes includes: filtering the brightness component and two chrominance components of the pixel points of the target image frame in three different neural network filtering processes respectively.

12. The method according to claim 11, characterized in that The filtering of the luminance component and two chrominance components of the pixel points of the target image frame in three different neural network filtering processes respectively includes: Performing a first neural network filtering process on a luminance component of a pixel point of the target image frame to perform filtering, and using the chrominance component as an input of the first neural network filtering process to filter the luminance component using the chrominance component; Performing a second neural network filtering process on a chrominance component of a pixel point of the target image frame to filter the chrominance component, and using the luminance component as an input of the second neural network filtering process to filter the one chrominance component using the other chrominance component; A third neural network filtering process is performed on the other chrominance component of the pixel point of the target image frame to filter it, and the luminance component is used as input of the third neural network filtering process to filter the other chrominance component using the one chrominance component.

13. The method according to any one of claims 1 to 6, characterized in that For each target image frame, filtering the pixel values ​​of the pixel points of the target image frame using the corresponding filtering processing method includes: Obtaining quantization parameters of a plurality of quantization blocks contained in a current frame corresponding to the target image frame; Neural network filtering corresponding to the preset range is performed on blocks corresponding to quantized blocks with quantization parameters in different preset ranges in the target image frame.

14. The method according to claim 13, wherein: For each of the quantization blocks, the step of performing neural network filtering corresponding to the preset range on blocks corresponding to the quantization blocks having quantization parameters in different preset ranges in the target image frame includes: Neural network filtering is performed on the luminance component and the chrominance component of the pixel points of the block corresponding to the quantized block in the target image frame.

15. The method according to claim 1, wherein Neural network filtering is performed using a neural network containing skip connections.

16. A method for encoding a video frame, characterized in that: include: The target image frame of the current frame is obtained by using the residual corresponding to the current frame; Performing a video frame filtering method on the target image frame to obtain a filtered target image frame, wherein the video frame filtering method is the method according to any one of claims 1 to 15; The current frame is encoded based on the filtered target image frame.

17. The method according to claim 16, characterized in that The method further includes: setting a first switch regarding a neural network filtering processing switch for a current sequence to which the current frame belongs; The method of filtering the video frame on the target image frame to obtain the filtered target image frame is performed when the first switch is on. The encoding of the current frame based on the filtered target image frame includes: Comparing the image quality of the filtered target image frame with that of the unfiltered target image frame; If the filtered target image frame is better than the unfiltered target image frame, setting the neural network filtering processing switch of the current frame to on; If the filtered target image frame is inferior to the unfiltered target image frame, the neural network filtering processing switch of the current frame is set to off.

18. The method according to claim 17, characterized in that The performing the video frame filtering method on the target image frame to obtain a filtered target image frame comprises: performing the video frame filtering method on the luminance component and two chrominance components of the target image frame respectively to obtain the filtered target image frame; The comparing the image quality of the filtered target image frame with that of the unfiltered target image frame comprises: comparing the image quality of each component of the filtered target image frame with that of the corresponding component of the unfiltered target image frame; If the filtered target image frame is better than the unfiltered target image frame, setting the neural network filtering processing switch of the current frame to on, including: if any component of the filtered target image frame is better than the corresponding component of the unfiltered target image frame, setting the neural network filtering processing switch corresponding to any component of the current frame to on; If the target image frame after filtering is inferior to the target image frame before filtering, the neural network filtering processing switch of the current frame is set to off, including: if any component of the target image frame after filtering is inferior to the corresponding component of the target image frame before filtering, the neural network filtering processing switch corresponding to any component of the current frame is set to off.

19. The method according to claim 16, wherein The performing the video frame filtering method according to any one of claims 1 to 15 on the target image frame to obtain a filtered target image frame comprises: performing the video frame filtering method according to any one of claims 1 to 15 on a plurality of target blocks obtained by segmenting the target image frame to obtain a filtered target image frame, The encoding of the current frame based on the filtered target image frame includes: Compare the image quality of each target block after filtering and the target block before filtering; If the filtered target block is better than the unfiltered target block, setting the neural network filtering processing switch corresponding to the target block to on; If the filtered target block is inferior to the unfiltered target block, setting the neural network filtering processing switch corresponding to the target block to off; Determining whether there is at least one target block whose corresponding neural network filter processing switch is set to on; If yes, setting the neural network filtering processing switch of the current frame to on; If not, the neural network filtering processing switch of the current frame is set to off.

20. A method for decoding a video frame, characterized in that: include: The target image frame of the current frame is obtained by using the residual corresponding to the current frame; The video frame filtering method according to any one of claims 1 to 15 is performed on the target image frame to obtain a filtered target image frame.

21. An encoder, characterized in that The encoder includes a processor and a memory coupled to each other, wherein the processor is configured to execute a computer program stored in the memory to implement the encoding method according to any one of claims 16 to 19.

22. A decoder, characterized in that The decoder includes a processor and a memory coupled to each other, wherein the processor is configured to execute a computer program stored in the memory to implement the decoding method according to claim 20.

23. A computer-readable storage medium, characterized in that A computer program that can be executed by a processor is stored, and the computer program is used to implement the filtering method of the video frame as described in any one of claims 1 to 15, or the encoding method as described in any one of claims 16 to 19, or the decoding method as described in claim 20.

Citation Information

Patent Citations

  • Filtering method and device

    CN110971915A

  • Image filtering method, device and equipment and storage medium

    CN111654710A

  • Loop filtering method, device and equipment in video encoding and decoding, and storage medium

    CN111711824A

  • Loop filtering method and device in video coding and decoding, equipment and storage medium

    CN113259671A