Filter network training method, video encoding method, device and electronic equipment

By filtering and encoding temporally correlated image frames and optimizing the filtering network with a differentiable encoder, the problem of high training costs of traditional filters and neural networks is solved, achieving efficient and accurate video encoding results.

CN116546224BActive Publication Date: 2026-03-10BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-23
Publication Date
2026-03-10

AI Technical Summary

Technical Problem

In existing technologies, traditional filters require manual adjustment before video encoding and have limited algorithm capabilities. Neural network-based filtering methods require a large number of pixel-level annotations, resulting in high training costs and insufficient filtering accuracy, which affects the overall video encoding effect.

Method used

By filtering multiple temporally correlated sample image frames based on an initial filtering network, and combining this with a differentiable encoder for image encoding and reconstruction, loss information is determined using image distortion information and bit rate, and the filtering network is optimized through unsupervised, end-to-end training.

Benefits of technology

It improves the training efficiency and accuracy of the filtering network, reduces the annotation cost, achieves higher video encoding quality, and eliminates the need for pixel-level label optimization of the filtering network.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116546224B_ABST
    Figure CN116546224B_ABST
Patent Text Reader

Abstract

The present disclosure relates to a training method of a filtering network, a video encoding method, a device and an electronic device. The method comprises: performing filtering processing on a plurality of sample image frames related in a time domain in a sample video based on an initial filtering network to obtain at least one sample filtered image corresponding to each sample image frame; inputting the at least one sample filtered image into a differentiable encoder to perform image encoding and reconstruction processing to obtain a plurality of sample reconstructed image frames corresponding to each sample image frame and a plurality of sample code rates corresponding to each sample image frame; determining image distortion information according to the plurality of sample reconstructed image frames and the plurality of sample image frames; determining loss information based on the image distortion information and the plurality of sample code rates; training the initial filtering network using the loss information until a loss condition is met, and using the initial filtering network corresponding to the time when the loss condition is met as a target filtering network. According to the technical solution provided by the present disclosure, the filtering accuracy of the target filtering network can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of artificial intelligence technology, and in particular to a training method for a filtering network, a video coding method, an apparatus, and an electronic device. Background Technology

[0002] With the rapid development of the internet, the amount of video data is growing exponentially. To cope with the storage and transmission pressure caused by this growth, video encoding and compression are chosen. In related technologies, traditional filters or neural networks are generally chosen for filtering before encoding. However, traditional filters have limited algorithmic capabilities and require manual adjustment, unable to optimize themselves. While neural network-based filtering can learn and optimize autonomously, current methods require a large number of labels (pixel-level labels) during training, and image filtering does not consider the impact of images on the overall video encoding, resulting in high labeling costs and inaccurate filtering learning leading to poor overall video encoding performance. Summary of the Invention

[0003] This disclosure provides a training method for a filtering network, a video coding method, an apparatus, and an electronic device to at least address the problem of how to improve the training efficiency and accuracy of filtering networks in related technologies. The technical solution of this disclosure is as follows:

[0004] According to a first aspect of the present disclosure, a method for training a filter network is provided, comprising:

[0005] Based on the initial filtering network, multiple sample image frames that are temporally correlated in the sample video are filtered to obtain at least one sample filtered image corresponding to each of the sample image frames.

[0006] At least one of the sample filtered images is input into a differentiable encoder for image encoding and reconstruction processing to obtain the sample reconstructed image frames corresponding to each of the plurality of sample image frames and their respective sample bit rates; the differentiable encoder refers to the function used by each module for encoding and reconstruction being differentiable or each module being a neural network;

[0007] Based on the reconstructed image frames from the multiple samples and the multiple sample image frames, the image distortion information is determined;

[0008] Based on the image distortion information and the bitrate of multiple samples, the loss information is determined;

[0009] The initial filtering network is trained using the loss information until the loss condition is met, and the initial filtering network corresponding to the condition is used as the target filtering network.

[0010] In one possible implementation, the plurality of sample image frames include a sample reference image frame and adjacent image frames of the sample reference image frame; the sample filtered image includes a first sample filtered image; the sample reconstructed image frame includes a first sample reconstructed image frame and a second sample reconstructed image frame; and the sample bitrate includes a first sample bitrate and a second sample bitrate.

[0011] The step of filtering multiple temporally correlated sample image frames in the sample video based on the initial filtering network to obtain at least one filtered sample image corresponding to each of the sample image frames includes:

[0012] The sample reference image frame and its adjacent image frames are input into the initial filtering network to filter the sample reference image frame, thereby obtaining the first sample filtered image corresponding to the sample reference image.

[0013] The step of inputting at least one of the sample filtered images into a differentiable encoder for image encoding and reconstruction processing to obtain the sample reconstructed image frames corresponding to each of the plurality of sample image frames and their respective sample bitrates includes:

[0014] The first sample filtered image is input into the differential encoder, and the first sample filtered image is encoded and reconstructed to obtain the first sample reconstructed image frame and the first sample bit rate.

[0015] The first sample reconstructed image frame and the adjacent image frames are input into the convertible encoder. The adjacent image frames are encoded and reconstructed to obtain the second sample reconstructed image frame and the second sample bitrate corresponding to each of the adjacent image frames.

[0016] In one possible implementation, the differentiable encoder includes a differentiable motion estimation network, a motion compensation network, a transform network, a quantization network, an inverse quantization network, an inverse transform network, a loop filtering network, and an entropy coding network.

[0017] The step of inputting the first sample filtered image into the differentiable encoder, encoding and reconstructing the first sample filtered image to obtain the first sample reconstructed image frame and the first sample bitrate includes:

[0018] Based on the motion estimation network, the motion compensation network, the transform network, and the quantization network, the first sample filtered image is sequentially subjected to motion estimation, motion compensation, transform encoding, and quantization processing to obtain image quantization data;

[0019] The image quantization data and the motion vector output by the motion estimation network are input into the entropy coding network for encoding processing to obtain the first sample bitrate.

[0020] The image quantization data is reconstructed using the inverse quantization network, the inverse transform network, the motion compensation network, and the loop filter network to obtain the first sample reconstructed image frame.

[0021] In one possible implementation, the step of inputting the sample reference image frame and its adjacent image frames into the initial filtering network to filter the sample reference image frame and obtain a first sample filtered image corresponding to the sample reference image includes:

[0022] Motion compensation processing is performed on the adjacent image frames to obtain the target adjacent image frames;

[0023] The sample reference image frame and the target adjacent image frame are input into the initial filtering network to filter the sample reference image frame, thereby obtaining the first sample filtered image corresponding to the sample reference image.

[0024] In one possible implementation, the sample filtered image includes a second sample filtered image; the step of filtering multiple temporally correlated sample image frames in the sample video based on the initial filtering network to obtain at least one sample filtered image corresponding to each of the sample image frames includes:

[0025] The plurality of sample image frames are input into the initial filtering network to perform filtering processing on the plurality of sample image frames, thereby obtaining the second sample filtered image corresponding to each of the plurality of sample image frames;

[0026] The step of inputting at least one of the sample filtered images into a differentiable encoder for image encoding and reconstruction processing to obtain the sample reconstructed image frames corresponding to each of the plurality of sample image frames and their respective sample bitrates includes:

[0027] Multiple second sample filtered images are input into the differential encoder, and the multiple second sample filtered images are encoded and reconstructed to obtain the multiple sample reconstructed image frames and the multiple sample bitrates.

[0028] In one possible implementation, the filtering process performed on multiple temporally correlated sample image frames in the sample video based on the initial filtering network to obtain at least one filtered sample image corresponding to each of the sample image frames includes:

[0029] Based on the initial filtering network, multiple temporally correlated sample image frames in the sample video are subjected to prediction processing of their respective filtering weights to obtain the sample filtering weights corresponding to each of the at least one sample image frame.

[0030] The at least one sample image frame is weighted according to the sample filtering weight to obtain the sample filtered image corresponding to each of the at least one sample image frame.

[0031] In one possible implementation, determining image distortion information based on the reconstructed image frames from the plurality of samples and the plurality of sample image frames includes:

[0032] Determine the image distortion sub-information between each sample image frame and the corresponding sample reconstructed image frame;

[0033] The image distortion information is obtained by weighting the multiple image distortion sub-information based on the distortion weights corresponding to the multiple sample image frames.

[0034] In one possible implementation, determining the loss information based on the image distortion information and multiple sample bitrates includes:

[0035] The sample bitrates are weighted based on the bitrate weights corresponding to the respective sample image frames to obtain the joint sample bitrate.

[0036] The image distortion information and the sample joint bitrate are summed to obtain the loss information.

[0037] In one possible implementation, the method further includes:

[0038] Acquire multiple sample images;

[0039] The multiple sample images are input into an initial differentiable encoder for image encoding and reconstruction processing to obtain multiple sample reconstructed images and multiple predicted code rates.

[0040] Based on the reconstructed image from the multiple samples, the multiple sample images, and the multiple predicted code rates, the coding loss information is determined;

[0041] The initial differentiable encoder is trained using the encoding loss information until the loss condition is met, and the initial differentiable encoder corresponding to the condition is taken as the differentiable encoder.

[0042] In one possible implementation, the resolution of the sample filtered image is lower than the resolution of the sample image frame; before determining the image distortion information based on the reconstructed image frames from the plurality of samples and the plurality of sample image frames, the method further includes:

[0043] Multiple reconstructed image frames from the samples are upsampled to obtain multiple upsampled images;

[0044] The step of reconstructing image frames based on multiple samples and determining image distortion information includes:

[0045] The image distortion information is determined based on the plurality of upsampled images and the plurality of sample image frames.

[0046] In one possible implementation, the upsampling process on the multiple reconstructed image frames to obtain multiple upsampled images includes:

[0047] Multiple reconstructed image frames from the samples are input into a preset upsampling network for upsampling processing to obtain the multiple upsampled images;

[0048] The method further includes: training the preset upsampling network using the loss information until the loss condition is met, and using the preset upsampling network corresponding to the loss condition as the target upsampling network.

[0049] According to a second aspect of the present disclosure, a video encoding method is provided, comprising:

[0050] Based on the target filtering network, multiple image frames to be encoded in the video to be encoded are filtered to obtain at least one target filtered image corresponding to each of the image frames to be encoded.

[0051] The target filtered image is input into the target encoder for image encoding processing to obtain the encoded image data corresponding to each of the multiple image frames to be encoded.

[0052] The target filtering network is trained according to the method described in any of the first aspects above.

[0053] In one possible implementation, the plurality of image frames to be encoded includes a target reference image frame and adjacent image frames that are temporally adjacent to the target reference image frame; the filtering process based on the target filtering network on the plurality of image frames to be encoded to obtain at least one target filtered image corresponding to each of the image frames to be encoded includes:

[0054] The target reference image frame and its temporally adjacent neighboring image frames are input into the target filtering network to filter the target reference image frame, thereby obtaining the first filtered image corresponding to the target reference image frame.

[0055] Accordingly, the step of inputting the target filtered image into the target encoder for image encoding processing to obtain the encoded image data corresponding to each of the plurality of image frames to be encoded includes:

[0056] The first filtered image and the target reference image frame are temporally adjacent adjacent image frames, which are input into the target encoder for image encoding processing to obtain the encoded image data corresponding to each of the plurality of image frames to be encoded.

[0057] In one possible implementation, before the step of filtering multiple image frames of the video to be encoded based on the target filtering network to obtain at least one target filtered image corresponding to each of the image frames to be encoded, the method further includes:

[0058] Obtain multiple initial image frames to be encoded from the video to be encoded;

[0059] The initial image frame is input to a filter for filtering to obtain the plurality of image frames to be encoded.

[0060] In one possible implementation, after the step of filtering multiple image frames of the video to be encoded based on the target filtering network to obtain at least one target filtered image corresponding to each of the image frames to be encoded, the method further includes:

[0061] The target filtered image is input into a filter for filtering to obtain a secondary filtered image;

[0062] The step of inputting the target filtered image into the target encoder for image encoding processing to obtain the encoded image data corresponding to each of the plurality of image frames to be encoded includes:

[0063] The secondary filtered image is input into the target encoder for image encoding processing to obtain the encoded image data corresponding to each of the multiple image frames to be encoded.

[0064] In one possible implementation, the resolution of the target filtered image is lower than the resolution of the image frame to be encoded; the method further includes:

[0065] The encoded image data is input into a target upsampling network for upsampling processing to obtain target encoded data; the target upsampling network is trained according to the method described in the last item of the first aspect above.

[0066] According to a third aspect of the present disclosure, a training apparatus for a filter network is provided, comprising:

[0067] The first filtering module is configured to perform filtering processing on multiple sample image frames that are temporally correlated in the sample video based on the initial filtering network, so as to obtain at least one sample filtered image corresponding to each of the sample image frames.

[0068] The encoding and reconstruction module is configured to input at least one of the sample filtered images into a differentiable encoder, perform image encoding and reconstruction processing, and obtain the sample reconstructed image frames corresponding to each of the plurality of sample image frames and their respective sample bit rates; the differentiable encoder refers to the function used by each module for encoding and reconstruction being differentiable or each module being a neural network;

[0069] The image distortion information determination module is configured to determine image distortion information based on the reconstructed image frames from the multiple samples and the multiple sample image frames.

[0070] The loss determination module is configured to determine loss information based on the image distortion information and multiple sample bitrates;

[0071] The training module is configured to train the initial filter network using the loss information until a loss condition is met, and then use the initial filter network corresponding to the condition as the target filter network.

[0072] In one possible implementation, the plurality of sample image frames include a sample reference image frame and adjacent image frames of the sample reference image frame; the sample filtered image includes a first sample filtered image; the sample reconstructed image frame includes a first sample reconstructed image frame and a second sample reconstructed image frame; and the sample bitrate includes a first sample bitrate and a second sample bitrate.

[0073] The first filtering module includes:

[0074] The first filtering unit is configured to input the sample reference image frame and its adjacent image frames into the initial filtering network, and perform filtering processing on the sample reference image frame to obtain the first sample filtered image corresponding to the sample reference image.

[0075] Accordingly, the encoding reconstruction module includes:

[0076] The reference frame coding and reconstruction unit is configured to input the first sample filtered image into the doubly encoder, encode and reconstruct the first sample filtered image to obtain a first sample reconstructed image frame and a first sample bitrate.

[0077] The adjacent frame encoding and reconstruction unit is configured to input the first sample reconstructed image frame and the adjacent image frame into the doubly encoder, and to encode and reconstruct the adjacent image frames to obtain the second sample reconstructed image frame and the second sample bitrate corresponding to each of the adjacent image frames.

[0078] In one possible implementation, the differentiable encoder includes a differentiable motion estimation network, a motion compensation network, a transform network, a quantization network, an inverse quantization network, an inverse transform network, a loop filtering network, and an entropy coding network.

[0079] The reference frame coding reconstruction unit includes:

[0080] The image quantization data acquisition subunit is configured to perform motion estimation, motion compensation, transformation encoding and quantization processing on the first sample filtered image in sequence based on the motion estimation network, the motion compensation network, the transform network and the quantization network to obtain image quantization data;

[0081] The first sample bitrate acquisition subunit is configured to perform encoding processing by inputting the image quantization data and the motion vector output by the motion estimation network into the entropy coding network to obtain the first sample bitrate.

[0082] The first sample reconstructed image frame acquisition subunit is configured to perform image reconstruction processing on the image quantization data using the inverse quantization network, the inverse transform network, the motion compensation network, and the loop filter network to obtain the first sample reconstructed image frame.

[0083] In one possible implementation, the first filtering unit includes:

[0084] The motion compensation subunit is configured to perform motion compensation processing on the adjacent image frames to obtain the target adjacent image frames;

[0085] The filtering subunit is configured to input the sample reference image frame and the target adjacent image frame into the initial filtering network, and perform filtering processing on the sample reference image frame to obtain the first sample filtered image corresponding to the sample reference image.

[0086] In one possible implementation, the first filtering module includes:

[0087] The second filtering unit is configured to input the plurality of sample image frames into the initial filtering network, perform filtering processing on the plurality of sample image frames, and obtain the second sample filtered image corresponding to each of the plurality of sample image frames.

[0088] Accordingly, the encoding reconstruction module includes:

[0089] The encoding and reconstruction unit is configured to input a plurality of second sample filtered images into the doubly encoder, encode and reconstruct the plurality of second sample filtered images to obtain the plurality of sample reconstructed image frames and the plurality of sample bitrates.

[0090] In one possible implementation, the first filtering module includes:

[0091] The filter weight prediction unit is configured to perform filter weight prediction processing on multiple sample image frames that are temporally correlated in the sample video based on the initial filter network, so as to obtain the sample filter weights corresponding to each of the at least one sample image frame.

[0092] The filtering unit is configured to perform weighted processing on the at least one sample image frame according to the sample filtering weight to obtain a sample filtered image corresponding to each of the at least one sample image frame.

[0093] In one possible implementation, the image distortion information determination module includes:

[0094] The image distortion sub-information determination unit is configured to determine the image distortion sub-information between each sample image frame and the corresponding sample reconstructed image frame.

[0095] The image distortion weighting unit is configured to perform weighting processing on multiple image distortion sub-information based on the distortion weights corresponding to the multiple sample image frames, so as to obtain the image distortion information.

[0096] In one possible implementation, the loss determination module includes:

[0097] The sample joint bitrate acquisition unit is configured to perform weighted processing on the bitrates of the multiple sample image frames based on the bitrate weights corresponding to each of the multiple sample image frames to obtain the sample joint bitrate.

[0098] The loss information acquisition unit is configured to perform summation processing on the image distortion information and the sample joint bit rate to obtain the loss information.

[0099] In one possible implementation, the device further includes:

[0100] The sample image acquisition module is configured to acquire multiple sample images;

[0101] The encoding prediction module is configured to perform image encoding and reconstruction processing by inputting the multiple sample images into an initial differentiable encoder, thereby obtaining multiple sample reconstructed images and multiple prediction code rates.

[0102] The coding loss information determination module is configured to determine coding loss information based on the reconstructed image from the plurality of samples, the plurality of sample images, and the plurality of predicted code rates.

[0103] The iterative training module of the differentiable encoder is configured to train the initial differentiable encoder using the encoding loss information until the loss condition is met, and then use the initial differentiable encoder corresponding to the loss condition as the differentiable encoder.

[0104] In one possible implementation, the resolution of the sample filtered image is lower than the resolution of the sample image frame; the apparatus further includes:

[0105] The first upsampling module is configured to perform upsampling processing on multiple reconstructed image frames of the samples to obtain multiple upsampled images;

[0106] The image distortion information determination module is further configured to determine the image distortion information based on the plurality of upsampled images and the plurality of sample image frames.

[0107] In one possible implementation, the first upsampling module includes:

[0108] The upsampling unit is configured to input multiple reconstructed image frames from the samples into a preset upsampling network for upsampling processing to obtain the multiple upsampled images.

[0109] The device further includes:

[0110] The upsampling network training module is configured to train the preset upsampling network using the loss information until the loss condition is met, and then use the preset upsampling network that meets the loss condition as the target upsampling network.

[0111] According to a fourth aspect of the present disclosure, a video encoding apparatus is provided, comprising:

[0112] The second filtering module is configured to perform filtering processing on multiple image frames to be encoded in the video to be encoded based on the target filtering network, so as to obtain at least one target filtered image corresponding to each of the image frames to be encoded.

[0113] The encoding module is configured to input the target filtered image into the target encoder for image encoding processing to obtain the encoded image data corresponding to each of the plurality of image frames to be encoded.

[0114] The target filtering network is trained according to the method described in the first aspect above.

[0115] In one possible implementation, the plurality of image frames to be encoded includes a target reference image frame and adjacent image frames that are temporally adjacent to the target reference image frame; the second filtering module includes:

[0116] The first filtered image acquisition unit is configured to input the target reference image frame and its temporally adjacent neighboring image frames into the target filtering network, perform filtering processing on the target reference image frame, and obtain the first filtered image corresponding to the target reference image frame.

[0117] Accordingly, the encoding module includes:

[0118] The first encoding unit is configured to input adjacent image frames that are temporally adjacent to the first filtered image and the target reference image frame into the target encoder for image encoding processing, thereby obtaining encoded image data corresponding to each of the plurality of image frames to be encoded.

[0119] In one possible implementation, the device further includes:

[0120] The initial image frame acquisition module is configured to acquire multiple initial image frames to be encoded in the video to be encoded;

[0121] The image frame acquisition module is configured to perform filtering processing on the initial image frame input filter to obtain the plurality of image frames to be encoded.

[0122] In one possible implementation, the device further includes:

[0123] The secondary filtering module is configured to perform filtering processing on the target filtered image input to the filter to obtain a secondary filtered image;

[0124] The encoding module includes:

[0125] The second encoding unit is configured to input the secondary filtered image into the target encoder for image encoding processing to obtain the encoded image data corresponding to each of the plurality of image frames to be encoded.

[0126] In one possible implementation, the resolution of the target filtered image is lower than the resolution of the image frame to be encoded; the apparatus further includes:

[0127] The second upsampling module is configured to perform upsampling by inputting the encoded image data into the target upsampling network for upsampling processing to obtain target encoded data; the target upsampling network is the target upsampling network trained in the third aspect above.

[0128] According to a fifth aspect of the present disclosure, an electronic device is provided, comprising: a processor; and a memory for storing processor-executable instructions; wherein the processor is configured to execute the instructions to implement the method as described in any one of the first aspects above.

[0129] According to a sixth aspect of the present disclosure, a computer-readable storage medium is provided such that, when instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to perform any of the methods described in the first aspect of the present disclosure.

[0130] According to a seventh aspect of the present disclosure, a computer program product is provided, including computer instructions that, when executed by a processor, cause a computer to perform the method described in any one of the first aspects of the present disclosure.

[0131] The technical solutions provided by the embodiments of this disclosure have at least the following beneficial effects:

[0132] By setting the input of the initial filtering network to multiple temporally correlated sample image frames, and combining a differentiable encoder to encode and reconstruct the filtered image, the reconstructed image frames corresponding to each of the multiple sample image frames and their corresponding sample bitrates are obtained. This allows the loss information to be obtained using the image distortion information from multiple frames and the multiple sample bitrates. This loss information effectively characterizes the filtering capability of the initial filtering network in subsequent encoding and reconstruction, thus better guiding the learning of the initial filtering network. The trained target filtering network can then perform filtering more efficiently and accurately, thereby improving the overall encoding efficiency and quality of the video. Furthermore, this method of training the initial filtering network using a differentiable encoder enables unsupervised, end-to-end optimization of the initial filtering network, eliminating the need for pixel-level labeling during the filtering process, saving resources, and further improving training efficiency.

[0133] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description

[0134] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure, and are not intended to unduly limit this disclosure.

[0135] Figure 1 This is a schematic diagram illustrating an application environment according to an exemplary embodiment.

[0136] Figure 2 This is a flowchart illustrating a training method for a filter network according to an exemplary embodiment.

[0137] Figure 3 This is a training process architecture diagram of a filtering network according to an exemplary embodiment.

[0138] Figure 4This is a schematic diagram of the structure of a directional encoder according to an exemplary embodiment.

[0139] Figure 5 This is a training process architecture diagram of another filtering network according to an exemplary embodiment.

[0140] Figure 6 This is a schematic diagram illustrating a video encoding process according to an exemplary embodiment.

[0141] Figure 7 This is a schematic diagram illustrating a two-stage filtering structure in video coding according to an exemplary embodiment.

[0142] Figure 8 This is a schematic diagram illustrating a structure of two filtering steps in another video coding method according to an exemplary embodiment.

[0143] Figure 9 This is a schematic diagram illustrating a joint training process of a filtering network and an upsampling network according to an exemplary embodiment.

[0144] Figure 10 This is a schematic diagram illustrating an encoding process under the condition of resolution change before and after filtering, according to an exemplary embodiment.

[0145] Figure 11 This is a block diagram of a training apparatus for a filter network according to an exemplary embodiment.

[0146] Figure 12 This is a block diagram illustrating an electronic device for video encoding according to an exemplary embodiment.

[0147] Figure 13 This is a block diagram illustrating an electronic device for training a filtering network according to an exemplary embodiment. Detailed Implementation

[0148] To enable those skilled in the art to better understand the technical solutions of this disclosure, the technical solutions in the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings.

[0149] It should be noted that the terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this disclosure are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this disclosure described herein can be implemented in orders other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this disclosure as detailed in the appended claims.

[0150] Artificial intelligence (AI) is the theory, methods, technology, and application systems that use digital computers or computers-controlled machines to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. AI software technology mainly includes computer vision, speech processing, natural language processing, and machine learning / deep learning.

[0151] In recent years, with the research and progress of artificial intelligence technology, artificial intelligence technology has been widely used in many fields. The solutions provided in the embodiments of this application involve technologies such as machine learning / deep learning, which are specifically illustrated through the following embodiments.

[0152] Please see Figure 1 , Figure 1 This is a schematic diagram illustrating an application environment according to an exemplary embodiment, such as... Figure 1 As shown, the application environment may include server 01 and terminal 02.

[0153] In an optional embodiment, server 01 can be used for training the filtering network. Specifically, server 01 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms.

[0154] In an optional embodiment, terminal 02 can be used for video encoding and transmission. Specifically, terminal 02 can be, but is not limited to, electronic devices such as smartphones, desktop computers, tablets, laptops, smart speakers, digital assistants, augmented reality (AR) / virtual reality (VR) devices, and smart wearable devices. Optionally, the operating system running on the electronic device can be, but is not limited to, Android, iOS, Linux, and Windows.

[0155] In addition, it should be noted that, Figure 1 The illustration shows only one application environment of the training method for the filtering network and the video coding method provided in this disclosure. Optionally, both the training method and the video coding method can be executed by server 01, and this disclosure does not limit this.

[0156] In the embodiments described in this specification, the server 01 and the terminal 02 can be directly or indirectly connected through wired or wireless communication, and this application does not impose any restrictions on this.

[0157] It should be noted that the following diagram illustrates one possible sequence of steps, and it is not strictly required to follow this order. Some steps can be performed in parallel without interdependence. The user information (including but not limited to user device information, user personal information, user behavior information, etc.) and data (including but not limited to data used for display, training data, etc.) involved in this disclosure are all information and data authorized by the user or fully authorized by all parties.

[0158] Before introducing the method embodiments provided in this application, a brief introduction will be given on the application scenarios, related terms or nouns that may be involved in the method embodiments of this application, so as to facilitate the understanding of those skilled in the art.

[0159] DNN (Deep Neural Network) is a type of deep neural network.

[0160] CNN (Convolutional Neural Network) is a type of neural network.

[0161] QP (Quantization Parameter) is the quantization coefficient.

[0162] MSE (Mean of Square Error).

[0163] SSIM (Structural Similarity Index) is a structural similarity index.

[0164] VMAF (Video Multimethod Assessment Fusion) is a multi-method fusion assessment of video quality.

[0165] MC (Motion Compensation).

[0166] ME (Motion Estimation).

[0167] MV (Motion Vector) is a vector representing motion.

[0168] DCT (Discrete Cosine Transform).

[0169] DST (The Discrete Cosine Transform) is an integer discrete sine transform.

[0170] HEVC (High Efficiency Video Coding) is a high-efficiency video coding system.

[0171] Figure 2 This is a flowchart illustrating a training method for a filter network according to an exemplary embodiment. Figure 2 As shown, it may include the following steps.

[0172] In step S201, multiple sample image frames that are temporally correlated in the sample video are filtered based on the initial filtering network to obtain at least one sample image frame corresponding to each sample image frame.

[0173] In the embodiments of this specification, the initial filtering network can refer to a deep neural network (DNN) used for training. The model parameters of the initial filtering network can be initialized. The initial filtering network, preset upsampling network, etc., mentioned in this disclosure can all refer to neural networks. The sample video can be multiple videos selected from a massive amount of video data; this disclosure does not limit the selection method. Multiple sample image frames that are temporally related can refer to temporally adjacent sample image frames; or they can refer to sample image frames from temporally adjacent sample image frames after removing redundant image frames. For example, in a sample video, image frames with frame numbers 1 to 8, where frames 6 and 7 have a large amount of redundancy, such as being in a static state, can be removed to form multiple sample image frames. Alternatively, they can be multiple sample image frames selected based on a sample reference image frame. For example, selecting a sample reference image frame and multiple adjacent image frames before and after it as multiple sample image frames. This disclosure does not limit this. As an example, the resolution of the sample filtered image can be the same as or different from the resolution of the sample image frame before filtering; this application does not limit this.

[0174] In step S203, at least one sample filtered image is input into a differentiable encoder for image encoding and reconstruction processing to obtain the sample reconstructed image frames corresponding to each of the multiple sample image frames and their corresponding sample bit rates. The differentiable encoder refers to the differentiable function used by each module for encoding and reconstruction or each module being a neural network, which enables subsequent gradient backpropagation based on loss information to adjust model parameters.

[0175] In one possible implementation, considering the relationship between the filtering and the reference frame in the encoding, the filtering is fused with the differentiable encoder. That is, the correlation between the frame encoding and the reference frame is considered during filtering. Therefore, the input can be set to include not only multiple frames but also reference frames. Specifically, multiple sample image frames can be selected based on sample reference image frames in the sample video. Based on this, multiple sample image frames can include the sample reference image frame and its adjacent image frames, where "adjacent" refers to temporal adjacency. Correspondingly, the sample filtered image can be the first sample filtered image, meaning that the sample filtered image corresponding to at least one sample image frame refers to the first sample filtered image corresponding to the sample reference image frame. The sample reconstructed image frame can include the first sample reconstructed image frame and the second sample reconstructed image frame; the sample bitrate can include the first sample bitrate and the second sample bitrate. The sample reference image frames in the sample video can be pre-defined, for example, image frames with frame numbers that are multiples of 8 can be pre-defined as sample reference image frames, such as frame 8, frame 16, frame 24, etc. Alternatively, the sample reference image frame in the sample video can be obtained based on preprocessing. For example, each image frame in the sample video can be coarsely encoded, allowing comparison of the encoding results and selection of the image frame with the better encoding result as the sample reference image frame. This application does not limit the method for selecting the reference image frame in the encoding of a video.

[0176] Accordingly, such as Figure 3 As shown, step S201 above may include: inputting the sample reference image frame and its adjacent image frames into the initial filtering network to filter the sample reference image frame, thereby obtaining the first sample filtered image corresponding to the sample reference image. That is, filtering the sample reference image frame based on temporally adjacent adjacent image frames, considering the temporal correlation between adjacent image frames during filtering, can effectively filter out information that is detrimental to subsequent encoding processing. The adjacent image frames of the sample reference image frame may include M frames preceding the sample reference image frame: I t-1 , ..., I t-M And N frames following the sample reference image frame: I t+1 , ..., I t+N N and M can be positive integers greater than or equal to 1, and the relationship between N and M can be the same or different; this disclosure does not impose any limitations on this. It should be noted that if the number of frames before the sample reference image frame is less than M or the number of frames after the sample reference image frame is less than N, then the number of adjacent image frames of the sample reference image frame can be less than N+M. That is, in this case, when acquiring adjacent image frames, all image frames before the sample reference image frame or all image frames after the sample reference image frame can be acquired.

[0177] When the input to the initial filtering network is a sample reference image frame and its adjacent image frames, step S203 can include: inputting the first sample filtered image into a differential encoder, encoding and reconstructing the first sample filtered image to obtain a first sample reconstructed image frame and a first sample bitrate. This first sample reconstructed image frame can be obtained by inversely decoding the quantized image data after transform encoding and quantization of the first sample filtered image. In other words, the differential encoder has both image encoding and image decoding (image reconstruction) functions.

[0178] Furthermore, the first sample reconstructed image frame and adjacent image frames can be input into a differentiable encoder to encode and reconstruct the adjacent image frames, obtaining the second sample reconstructed image frames and second sample bitrates corresponding to each adjacent image frame. See also Figure 3 The dashed section indicates that after obtaining the first sample reconstructed image frame, the first sample reconstructed image frame and each adjacent image frame can be sequentially input into the differential encoder. Each adjacent image frame is then encoded and reconstructed sequentially to obtain the second sample reconstructed image frame and the second sample bitrate corresponding to each adjacent image frame. In this way, the first sample bitrate and the second sample bitrate can form multiple sample bitrates; the first sample reconstructed image frame and the second sample reconstructed image frame can form multiple sample reconstructed image frames.

[0179] Optionally, before performing step S203, the sample filtered image and the sample image frame can be fused to obtain a fused filtered image. Correspondingly, step S203 can be replaced by: inputting the fused filtered image into a differentiable encoder for image encoding and reconstruction processing to obtain multiple reconstructed sample images and their corresponding sample bitrates. That is, before inputting into the differentiable encoder, the sample filtered image and the corresponding sample image frame can be fused. For cases with sample reference image frames, such fusion processing can also be performed accordingly, which will not be elaborated further here. As an example, the fusion processing here can refer to image weighting processing, which is not limited in this disclosure.

[0180] By setting sample reference image frames and filtering them during the filtering stage, it is not necessary to filter all image frames, thus reducing the number of image frames to be filtered. Furthermore, the filtering of sample reference image frames considers adjacent image frames in the temporal domain, making the filtering of sample image frames more in line with the needs of subsequent encoding. This achieves effective integration of filtering and encoding, improving both encoding efficiency and encoding quality. In other words, by using sample reference image frames and their adjacent image frames as training samples, the trained target filtering network can perform image filtering more accurately and more effectively serve the overall video encoding, improving the overall efficiency and quality of video encoding.

[0181] Reference Figure 4 In one example, the modules used for encoding and reconstruction in a differentiable encoder can be neural networks. In this case, the structure of the differentiable encoder can be as follows: Figure 4 As shown, a differentiable encoder may include a differentiable motion estimation network (ME), a motion compensation network (MC), a transform network, a quantization network, an inverse quantization network, an inverse transform network, a loop filter network, and an entropy coding network. Optionally, as... Figure 4 As shown, it may also include a decoded image frame buffer module for storing decoded image frames. Correspondingly, the step of inputting the first sample filtered image into a differentiable encoder and performing encoding and reconstruction processing on the first sample filtered image to obtain the first sample reconstructed image frame and the first sample bitrate may include: sequentially performing motion estimation, motion compensation, transformation encoding, and quantization processing on the first sample filtered image based on a motion estimation network, a motion compensation network, a transform network, and a quantization network to obtain image quantization data; thereby, the image quantization data and the motion vector output by the motion estimation network can be input into an entropy encoding network for encoding processing to obtain the first sample bitrate. Furthermore, the image quantization data can be reconstructed using an inverse quantization network, an inverse transform network, a motion compensation network, and a loop filtering network to obtain the first sample reconstructed image frame.

[0182] Specifically, see Figure 4 The first sample filtered image can be input into the motion estimation network ME for motion estimation, obtaining the motion vector MV. MV is then used as input to MC for motion compensation processing to obtain the predicted image. The difference between the first sample filtered image and the predicted image is then calculated to obtain a residual image. A transform network is then used to perform transform coding on the residual image to obtain transform-coded data. Further, the transform-coded data can be input into a quantization network and quantized using quantization coefficients QP to obtain quantized transform-coded data, i.e., image quantization data. Based on this, the image quantization data and MV can be input into an entropy coding network for entropy coding to obtain the first sample bitrate.

[0183] Accordingly, inverse operations such as inverse quantization and inverse transform can be performed on the quantized transform-coded data to obtain the inverse transform residual image. The predicted image output from the MV processing based on MC and the inverse transform residual image can be summed to obtain the initial reconstructed image. Based on this, the initial reconstructed image can be input into a loop filter network for filtering to obtain the first sample reconstructed image frame.

[0184] It should be noted that, Figure 4 The decoding indicated by the dashed line can refer to the structure used for decoding in a differentiable encoder. The MC (Metal-Chip) can be considered a network shared by both encoding and decoding; that is, decoding yields the MV (Multi-Image), so the MC is needed to process it to obtain the predicted image. Optionally, decoding can also yield the predicted image output by the MC. Based on this... Figure 4 The schematic diagram of the decoding structure indicated by the dashed line does not limit this disclosure. Based on the above description, the reconstruction process can be implemented based on the decoding structure; therefore, a differentiable encoder can be referred to as a differentiable codec.

[0185] The first sample reconstructed image frame can be stored in the decoded image frame buffer module and can be used as the output of the differentiable encoder. By using a differentiable function or neural network to approximate the traditional encoder, the encoder becomes differentiable, which enables gradient backpropagation to autonomously optimize the parameters of the filtering network. Furthermore, the differentiable encoder in this disclosure has a simple structure, making the encoding process more efficient.

[0186] As an example, the modules in a differentiable encoder can be configured based on the functions of the modules in HEVC. For instance, the motion compensation (MC) can utilize a neural network to implement an 8-tap or 4-tap interpolation filter consistent with HEVC, making the MC differentiable. The differentiable transformation functions of the transform and inverse transform networks can be block-level DCT or DST transformation functions in HEVC. The quantization network can perform uniform scalar quantization on all transform coefficients according to quantization parameters, consistent with scalar quantization in HEVC. The loop filtering network can use a CNN network for filtering. The entropy coding network can use a context model to probabilistically model and entropy-encode all quantized transform coefficients, similar to the context-based adaptive binary arithmetic coding algorithm in HEVC. This disclosure does not limit the specific differentiable functions or neural networks of the modules in the differentiable encoder.

[0187] In an optional implementation, the aforementioned differentiable encoder can be pre-trained and its parameters fixed before being used to train the filtering network. Based on this, the training process of the differentiable encoder can include the following steps: acquiring multiple sample images, for example, images from training sample videos can be acquired as multiple sample images; further, the multiple sample images can be input into the initial differentiable encoder for image encoding and reconstruction processing to obtain multiple reconstructed sample images and multiple predicted bitrates; determining encoding loss information based on the multiple reconstructed sample images, the multiple sample images, and the multiple predicted bitrates; thereby, the initial differentiable encoder can be trained using the encoding loss information until a loss condition is met, and the initial differentiable encoder corresponding to the met loss condition is taken as the differentiable encoder. The loss condition can be a loss threshold, which is not limited in this disclosure.

[0188] Specifically, training an initial differentiable encoder using encoding loss information can include calculating gradients using the encoding loss information and adjusting the parameters of each network in the differentiable encoder using gradient backpropagation; this process is iterated until the loss information satisfies the loss condition, and the initial differentiable encoder that satisfies the loss condition is taken as the differentiable encoder.

[0189] It should be noted that the processing and output of the input sample images during the training of the differentiable encoder can be found above. Figure 4 The relevant information will not be repeated here.

[0190] During the training of a differentiable encoder, for example, if the input sample image is I, the output sample reconstructed image after passing through the differentiable encoder is... The distortion function is L(), and the bit rate of the differentiable encoder output is R. I That is, the predicted bitrate can be achieved using R I This indicates that a hyperparameter λ1 can be given to control the bit rate point of the differentiable encoder, optimizing the model parameters of the entire differentiable encoder to minimize the following loss function:

[0191]

[0192] Here, `loss` represents the encoding loss information; the function `L()` can be one of the quality assessment functions such as MSE, SSIM, and VMAF; `λ1` can be set based on compression requirements. The higher the compression requirements, i.e., the lower the compression ratio, which means the smaller the compressed file, the larger `λ1` can be. In one example, the value of `λ1` can be from 10 to 200.

[0193] Optionally, after training the differentiable encoder, the parameters of the differentiable encoder model are fixed for use in training subsequent filtering networks.

[0194] In one alternative implementation, refer to Figure 5Motion compensation can be applied to adjacent image frames before inputting them into the initial filtering network, which can improve the filtering efficiency and accuracy of the sample reference image frame. Based on this, the above-mentioned inputting the sample reference image frame and its adjacent image frames into the initial filtering network to filter the sample reference image frame and obtain the first sample filtered image corresponding to the sample reference image can include: performing motion compensation processing on adjacent image frames to obtain the target adjacent image frame; thus, the sample reference image frame and the target adjacent image frame can be input into the initial filtering network to filter the sample reference image frame and obtain the first sample filtered image corresponding to the sample reference image.

[0195] Reference Figure 5 As shown, for the current frame I to be filtered in the training data... t (i.e., the sample reference image frame in the sample video), can be compared with the current frame I. t Adjacent preceding M frames I t-1 , ..., I t-M and the adjacent N frames I t+1 , ..., I t+N Perform motion compensation, then add the current frame I t The M+N frames (i.e., the target's adjacent image frames) after motion compensation are input into the initial filtering network to be trained. The temporal filtering result I′ of the current frame is inferred from the initial filtering network. t , that is, the first sample filtered image;

[0196] Furthermore, the time-domain filtering result I′ can be... t The reconstructed frame is obtained by inputting the pre-trained differentiable encoder. and encoding rate R t That is, the first sample reconstructed image frame and the first sample bitrate are obtained.

[0197] Then, the reconstructed frame of the current frame can be... As adjacent frame I t-1 , ..., I t-M I t+1 , ..., I t+N The reference frame is used for encoding, and adjacent frames are encoded sequentially using a pre-trained differentiable encoder.

[0198] I t-1 , ..., I t-M I t+1 , ..., I t+N The corresponding reconstructed frame is obtained by encoding. and encoding rate R t-1 , ..., R t-M R t+1 , ..., R t+M That is, to obtain the second sample reconstructed image frame and the second sample bitrate.

[0199] In step S205, image distortion information is determined based on multiple sample reconstructed image frames and multiple sample image frames.

[0200] In the embodiments of this specification, image distortion information can be determined based on the distortion of multiple frames of images. For example, the average distortion of multiple sample reconstructed image frames and multiple sample image frames can be determined as image distortion information.

[0201] In one possible implementation, determining image distortion information based on multiple sample reconstructed image frames and multiple sample image frames may include: determining image distortion sub-information between each sample image frame and its corresponding sample reconstructed image frame; and weighting the multiple image distortion sub-information based on the distortion weights corresponding to each of the multiple sample image frames to obtain the image distortion information. The image distortion information can be represented by D, where D is the standard value of D. Figure 5 For example, taking the above I t Taking N+M adjacent image frames as an example, the specific calculation of D can be as follows:

[0202]

[0203] Where, ω i Let I be the weight of the distortion in the i-th sample image frame relative to the loss information (Loss), i.e., the distortion weight; i For the i-th sample image frame; The image frame to be reconstructed is the image frame corresponding to the i-th sample image frame, i.e., the image frame to be reconstructed from the i-th sample; L() is a function to calculate distortion, such as one of the functions MSE, SSIM, VMAF, etc. This represents the image distortion sub-information between the i-th sample image frame and the corresponding reconstructed sample image frame.

[0204] In step S207, loss information is determined based on image distortion information and multiple sample bit rates.

[0205] In the embodiments of this specification, a preset rate-distortion loss function can be used to calculate the loss of image distortion information and multiple sample bitrates to obtain loss information. This disclosure does not limit the preset rate-distortion loss function.

[0206] In one possible implementation, determining the loss information based on image distortion information and multiple sample bitrates can include: weighting multiple sample bitrates based on the bitrate weights corresponding to each of the multiple sample image frames to obtain a joint sample bitrate; thereby, the image distortion information and the joint sample bitrate can be summed to obtain the loss information.

[0207] by Figure 5 For example, the sample joint code rate R can be calculated using the following formula:

[0208]

[0209] Where, α i R represents the weight of the sample bitrate of the i-th sample image frame in relation to the loss information (Loss); i Let be the sample bitrate of the i-th sample image frame. The sample bitrate may include the first sample bitrate and the second sample bitrate.

[0210] Furthermore, the loss information (Loss) can be calculated using the following formula:

[0211] Loss=D+λR

[0212] Here, λ can be a hyperparameter controlling the bit rate point. This λ can be set based on compression requirements. The higher the compression requirements, i.e., the lower the compression ratio, which means the smaller the compressed file, the larger λ can be. In one example, the value of λ can be from 10 to 200.

[0213] In step S209, the initial filtering network is trained using the loss information until the loss condition is met, and the initial filtering network corresponding to the loss condition is taken as the target filtering network.

[0214] It should be noted that the parameters of the differentiable encoder are fixed during the training of the initial filtering network. That is, the parameters of the pre-trained differentiable encoder are not changed during the training of the initial filtering network. The model parameters of the initial filtering network are optimized only by minimizing the multi-frame joint rate-distortion cost loss function Loss (i.e., the aforementioned loss information Loss) using the gradient descent method. This achieves unsupervised end-to-end optimization of the initial filtering network in the encoding process. The training to optimize the model parameters of the initial filtering network can be done by calculating gradients based on the loss information, and then using gradient descent for gradient backpropagation to adjust the model parameters of the initial filtering network, completing one iteration. This process is repeated until the loss information meets the loss condition. The initial filtering network corresponding to the loss condition can then be used as the target filtering network. The initial filtering network with updated model parameters from the previous iteration is used in the next iteration.

[0215] By setting the input of the initial filtering network to multiple temporally correlated sample image frames, and combining it with a differentiable encoder to encode and reconstruct the filtered image, the corresponding sample reconstructed image frames and their corresponding sample bitrates are obtained. This allows the loss information to be obtained using the image distortion information of multiple frames and the multiple sample bitrates. The loss information effectively characterizes the filtering capability of the initial filtering network in subsequent encoding and reconstruction, thus better guiding the learning of the initial filtering network. This enables the trained target filtering network to perform filtering more efficiently and accurately, thereby improving the overall encoding efficiency and quality of the video. For example, when filtering a single frame input, some high-frequency textures and other noises might be directly filtered out. However, using the scheme disclosed herein, since the correlation between temporally adjacent frames is considered during filtering, high-frequency textures and other noises will be retained when they are also present in adjacent frames, without denoising. This is more beneficial for subsequent encoding processing, improving encoding efficiency and encoding effect.

[0216] In addition, this method of training the initial filter network using a differentiable encoder can achieve unsupervised, end-to-end optimization of the initial filter network without the need to label pixels during the filtering process, saving resources and further improving training efficiency.

[0217] In another possible implementation, the initial filtering network can be a multi-input, multi-output (MIMO) network, meaning that the multiple sample image frames do not include the sample reference image frame. Based on this, the sample filtered image can include a second sample filtered image. Accordingly, filtering multiple temporally correlated sample image frames in the sample video based on the initial filtering network to obtain at least one sample filtered image corresponding to each sample image frame can include: inputting multiple sample image frames into the initial filtering network, filtering the multiple sample image frames, and obtaining a second sample filtered image corresponding to each of the multiple sample image frames.

[0218] Building upon the aforementioned multiple-input multiple-output (MIMO) approach, at least one sample filtered image is input into a differentiable encoder for image encoding and reconstruction, yielding multiple sample image frames and their corresponding reconstructed image frames and bitrates. This can include inputting multiple second sample filtered images into the differentiable encoder and encoding them to obtain multiple reconstructed image frames and bitrates. The encoding and reconstruction of each second sample filtered image in the differentiable encoder can refer to the processing of the sample reference image frame input into the differentiable encoder; that is, the encoding and reconstruction process of the reference image frame is not required, and will not be elaborated further here. This MIMO approach also considers the temporal correlation of image frames, improving filtering efficiency and effectiveness.

[0219] The initial filtering network described above outputs the filtered pixel result. In another possible implementation, the output of the initial filtering network can also be filtering weights, which are then used to weight the sample images to obtain the filtered image. Based on this, the above-mentioned filtering process performed on multiple temporally correlated sample image frames in the sample video using the initial filtering network to obtain at least one sample filtered image corresponding to each of the sample image frames can include: predicting the filtering weights of multiple temporally correlated sample image frames in the sample video using the initial filtering network to obtain the sample filtering weights corresponding to each of the at least one sample image frame; and weighting the at least one sample image frame according to the sample filtering weights to obtain the sample filtered image corresponding to each of the at least one sample image frame. Here, the at least one sample image frame can be determined according to the above-mentioned multiple-input multiple-output or multiple-input one-output method. Multiple-input one-output refers to the case where there is a sample reference image frame. In this case, the sample filtering weights corresponding to each of the at least one sample image frame can refer to the sample filtering weights corresponding to the sample reference image frame, and correspondingly, the first sample filtered image corresponding to the sample reference image frame can be obtained. Multiple-input multiple-output (MIMO) refers to the case where no specific sample reference image frame is specified. In this case, the sample filtering weights corresponding to at least one sample image frame can refer to the sample filtering weights corresponding to multiple sample image frames. Accordingly, second sample filtered images corresponding to multiple sample image frames can be obtained. By learning the filtering weights, training efficiency can be improved.

[0220] After obtaining the target filter network through the above training, it can be applied online, that is, video encoding can be performed using the target filter network obtained by the above training method. Based on this, this disclosure also provides a video encoding method, which may include:

[0221] The target filtering network filters multiple image frames of the video to be encoded, obtaining at least one target filtered image corresponding to each image frame. The processing procedure for this step is similar to that in step S201 above and will not be repeated here. The target filtering network can be trained on an initial filtering network based on image rate-distortion loss information. This loss information is determined by encoding and reconstructing sample images, sample bitrates, and sample image frames obtained by a differentiable encoder from multiple temporally correlated sample image frames input to the sample filtered images of the initial filtering network.

[0222] Furthermore, the target filtered image can be input into the target encoder for image encoding processing to obtain encoded image data corresponding to each of the multiple image frames to be encoded. The target encoder may differ from the differentiable encoder; it can be a traditional encoder used for video encoding, and the modules of the target encoder do not need to be differentiable.

[0223] By using the target filtering network obtained through the above training method for filtering and subsequent encoding, video encoding efficiency and quality can be improved.

[0224] In the embodiments of this specification, when applied online, the input and output methods of the target filtering network can also include multiple-input one-output (MIMO) and multiple-input multiple-output (MIMO) methods. For specific processing flows, please refer to the relevant description in the training section above, which will not be repeated here. Taking multiple-input multiple-output as an example, the multiple image frames to be encoded can include a target reference image frame and adjacent image frames that are temporally adjacent to the target reference image frame. Accordingly, filtering the multiple image frames to be encoded from the video to be encoded based on the target filtering network to obtain at least one target filtered image corresponding to each of the image frames to be encoded can include:

[0225] The target reference image frame and its temporally adjacent adjacent image frames are input into the target filtering network to filter the target reference image frame, thereby obtaining the first filtered image corresponding to the target reference image frame.

[0226] Accordingly, the target filtered image is input into the target encoder for image encoding processing to obtain encoded image data corresponding to each of the multiple image frames to be encoded, which may include:

[0227] The first filtered image and the target reference image frame are temporally adjacent adjacent image frames, which are input into the target encoder for image encoding processing to obtain the encoded image data, such as bitstreams, corresponding to each of the multiple image frames to be encoded.

[0228] By specifying a target reference image frame and combining it with adjacent image frames in the temporal domain, the target reference image frame is filtered. This allows the filtering of the target reference image frame to more effectively serve subsequent coding processes, which not only reduces the number of frames to be filtered but also improves the efficiency and quality of subsequent coding.

[0229] Optionally, in the case of multiple inputs and multiple outputs, multiple image frames of the video to be encoded can be input into the target filtering network for filtering processing to obtain the second filtered image corresponding to each of the multiple image frames to be encoded.

[0230] In an optional implementation, if the output of the trained target filtering network is filter weights, then filtering is performed on multiple image frames to be encoded in the video to be encoded based on the target filtering network to obtain at least one target filtered image corresponding to each image frame. This can include: predicting the filter weights of each of the multiple image frames to be encoded based on the target filtering network to obtain at least one target filter weight corresponding to each image frame; thereby, at least one image frame to be encoded can be weighted according to the target filter weights to obtain at least one target filtered image corresponding to each image frame. For a detailed description of the sample filter weights, please refer to the above introduction, which will not be repeated here.

[0231] Optionally, in video encoding applications, the target filtered image and the image frame to be encoded can be fused before being input into the target encoder for image encoding. Specifically, the target filtered image and the image frame to be encoded can be fused, for example, through image weighting, to obtain a fused image. Then, the fused image can be input into the target encoder for image encoding to obtain encoded image data corresponding to each of the multiple image frames to be encoded.

[0232] See Figure 7 and Figure 8 In practical applications, the method can use only the target filtering network as described above, or it can be combined with a traditional filter for two filtering processes. Therefore, before the step of filtering multiple image frames of the video to be encoded based on the target filtering network to obtain at least one target filtered image corresponding to each image frame, the method further includes: acquiring multiple initial image frames to be encoded in the video; and inputting the initial image frames into a filter for filtering to obtain multiple image frames to be encoded.

[0233] Optionally, after the step of filtering multiple image frames of the video to be encoded based on the target filtering network to obtain at least one target filtered image corresponding to each of the image frames to be encoded, the method may further include: inputting the target filtered image into a filter for filtering to obtain a secondary filtered image.

[0234] Accordingly, the above-mentioned inputting the target filtered image into the target encoder for image encoding processing to obtain the encoded image data corresponding to each of the multiple image frames to be encoded includes: inputting the secondary filtered image into the target encoder for image encoding processing to obtain the encoded image data corresponding to each of the multiple image frames to be encoded.

[0235] By combining the target filtering network with a traditional filter, two filtering processes are performed, making the filtering more accurate and improving the efficiency of subsequent coding.

[0236] In an optional implementation, the resolution of the images before and after filtering by the initial filtering network is variable. For example, the resolution of the sample filtered image can be lower than the resolution of the sample image frame, thus adapting to the compression requirements of the differential encoder. That is, when the image resolution before and after filtering is variable, the initial filtering network can be regarded as an initial downsampling network. Based on this, before determining the image distortion information based on multiple sample reconstructed image frames and multiple sample image frames, the training method may further include: performing upsampling processing on the multiple sample reconstructed image frames to obtain multiple upsampled images. The resolution of the upsampled image can be the same as the resolution of the corresponding sample image frame.

[0237] Accordingly, the above-mentioned determination of image distortion information based on reconstructed image frames from multiple samples and multiple sample image frames can be replaced by: determining image distortion information based on multiple upsampled images and multiple sample image frames. The specific determination process can be found in step S205 above, and will not be repeated here. Further, subsequent steps S207 to S209 can be performed to train the initial filtering network.

[0238] By setting the initial filter network to have a lower resolution after filtering, the amount of data processed by the differentiable encoder can be reduced, thereby improving the training efficiency of the initial filter network.

[0239] Optionally, see Figure 9 Upsampling can be achieved through neural networks. Therefore, the above-mentioned upsampling process for multiple reconstructed image frames to obtain multiple upsampled images can include: inputting multiple reconstructed image frames into a preset upsampling network for upsampling processing to obtain multiple upsampled images. Correspondingly, the training method can also include: training the preset upsampling network using loss information until a loss condition is met, thus using the preset upsampling network that meets the loss condition as the target upsampling network. In this joint training of the initial filtering network and the preset upsampling network, the model parameters of the initial filtering network and the preset upsampling network are updated according to the loss information in each iteration. In the next training iteration, the updated initial filtering network and the preset upsampling network are used to repeat the above steps to obtain new loss information for updating the model parameters of the initial filtering network and the preset upsampling network in the next training iteration. This iterative process is repeated until the loss information meets the loss condition. This enables joint training of the initial upsampling network and the preset downsampling network, which can improve training efficiency and the encoding effect under subsequent image resolution changes.

[0240] Reference Figure 10Given the changes in image resolution before and after filtering, in video coding applications, the resolution of the target filtered image is lower than the resolution of the frame to be encoded. The video coding method may further include: inputting the encoded image data into a target upsampling network for upsampling processing to obtain target encoded data; the target upsampling network is obtained according to the training method described above. Since the target filtering network can reduce the resolution of the input image, the amount of data processed by the target encoder is reduced, improving coding efficiency. Combined with the target upsampling network, the resolution of the image to be encoded can be maintained, thus simultaneously improving both coding effect and coding efficiency.

[0241] Figure 11 This is a block diagram illustrating a training apparatus for a filter network according to an exemplary embodiment. (Refer to...) Figure 11 The device may include:

[0242] The first filtering module 1101 is configured to perform filtering processing on multiple sample image frames that are temporally correlated in the sample video based on the initial filtering network, so as to obtain at least one sample filtered image corresponding to each of the sample image frames.

[0243] The encoding and reconstruction module 1103 is configured to input at least one of the sample filtered images into a differentiable encoder, perform image encoding and reconstruction processing, and obtain the sample reconstructed image frames corresponding to each of the plurality of sample image frames and their respective sample bit rates; the differentiable encoder refers to the function used by each module for encoding and reconstruction being differentiable or each module being a neural network;

[0244] The image distortion information determination module 1105 is configured to perform image distortion information determination based on the reconstructed image frames from the plurality of samples and the plurality of sample image frames.

[0245] The loss determination module 1107 is configured to determine loss information based on the image distortion information and multiple sample bitrates;

[0246] Training module 1109 is configured to train the initial filter network using the loss information until the loss condition is met, and to use the initial filter network corresponding to the loss condition as the target filter network.

[0247] By setting the input of the initial filtering network to multiple temporally correlated sample image frames, and combining a differentiable encoder to encode and reconstruct the filtered image, the reconstructed image frames corresponding to each of the multiple sample image frames and their corresponding sample bitrates are obtained. This allows the loss information to be obtained using the image distortion information from multiple frames and the multiple sample bitrates. This loss information effectively characterizes the filtering capability of the initial filtering network in subsequent encoding and reconstruction, thus better guiding the learning of the initial filtering network. The trained target filtering network can then perform filtering more efficiently and accurately, thereby improving the overall encoding efficiency and quality of the video. Furthermore, this method of training the initial filtering network using a differentiable encoder enables unsupervised, end-to-end optimization of the initial filtering network, eliminating the need for pixel-level labeling during the filtering process, saving resources, and further improving training efficiency.

[0248] In one possible implementation, the plurality of sample image frames include a sample reference image frame and adjacent image frames of the sample reference image frame; the sample filtered image includes a first sample filtered image; the sample reconstructed image frame includes a first sample reconstructed image frame and a second sample reconstructed image frame; and the sample bitrate includes a first sample bitrate and a second sample bitrate.

[0249] The first filtering module 1101 may include:

[0250] The first filtering unit is configured to input the sample reference image frame and its adjacent image frames into the initial filtering network, and perform filtering processing on the sample reference image frame to obtain the first sample filtered image corresponding to the sample reference image.

[0251] Accordingly, the encoding reconstruction module 1103 may include:

[0252] The reference frame coding and reconstruction unit is configured to input the first sample filtered image into the doubly encoder, encode and reconstruct the first sample filtered image to obtain a first sample reconstructed image frame and a first sample bitrate.

[0253] The adjacent frame encoding and reconstruction unit is configured to input the first sample reconstructed image frame and the adjacent image frame into the doubly encoder, and to encode and reconstruct the adjacent image frames to obtain the second sample reconstructed image frame and the second sample bitrate corresponding to each of the adjacent image frames.

[0254] In one possible implementation, the differentiable encoder includes a differentiable motion estimation network, a motion compensation network, a transform network, a quantization network, an inverse quantization network, an inverse transform network, a loop filtering network, and an entropy coding network.

[0255] The reference frame coding reconstruction unit includes:

[0256] The image quantization data acquisition subunit is configured to perform motion estimation, motion compensation, transformation encoding and quantization processing on the first sample filtered image in sequence based on the motion estimation network, the motion compensation network, the transform network and the quantization network to obtain image quantization data;

[0257] The first sample bitrate acquisition subunit is configured to perform encoding processing by inputting the image quantization data and the motion vector output by the motion estimation network into the entropy coding network to obtain the first sample bitrate.

[0258] The first sample reconstructed image frame acquisition subunit is configured to perform image reconstruction processing on the image quantization data using the inverse quantization network, the inverse transform network, the motion compensation network, and the loop filter network to obtain the first sample reconstructed image frame.

[0259] In one possible implementation, the first filtering unit includes:

[0260] The motion compensation subunit is configured to perform motion compensation processing on the adjacent image frames to obtain the target adjacent image frames;

[0261] The filtering subunit is configured to input the sample reference image frame and the target adjacent image frame into the initial filtering network, and perform filtering processing on the sample reference image frame to obtain the first sample filtered image corresponding to the sample reference image.

[0262] In one possible implementation, the first filtering module 1101 may include:

[0263] The second filtering unit is configured to input the plurality of sample image frames into the initial filtering network, perform filtering processing on the plurality of sample image frames, and obtain the second sample filtered image corresponding to each of the plurality of sample image frames.

[0264] Accordingly, the encoding reconstruction module 1103 may include:

[0265] The encoding and reconstruction unit is configured to input a plurality of second sample filtered images into the doubly encoder, encode and reconstruct the plurality of second sample filtered images to obtain the plurality of sample reconstructed image frames and the plurality of sample bitrates.

[0266] In one possible implementation, the first filtering module 1101 may include:

[0267] The filter weight prediction unit is configured to perform filter weight prediction processing on multiple sample image frames that are temporally correlated in the sample video based on the initial filter network, so as to obtain the sample filter weights corresponding to each of the at least one sample image frame.

[0268] The filtering unit is configured to perform weighted processing on the at least one sample image frame according to the filtering weight to obtain a sample filtered image corresponding to each of the at least one sample image frame.

[0269] In one possible implementation, the image distortion information determination module 1105 may include:

[0270] The image distortion sub-information determination unit is configured to determine the image distortion sub-information between each sample image frame and the corresponding sample reconstructed image frame.

[0271] The image distortion weighting unit is configured to perform weighting processing on multiple image distortion sub-information based on the distortion weights corresponding to the multiple sample image frames, so as to obtain the image distortion information.

[0272] In one possible implementation, the loss determination module 1107 may include:

[0273] The sample joint bitrate acquisition unit is configured to perform weighted processing on the bitrates of the multiple sample image frames based on the bitrate weights corresponding to each of the multiple sample image frames to obtain the sample joint bitrate.

[0274] The loss information acquisition unit is configured to perform summation processing on the image distortion information and the sample joint bit rate to obtain the loss information.

[0275] In one possible implementation, the device may further include:

[0276] The sample image acquisition module is configured to acquire multiple sample images;

[0277] The encoding prediction module is configured to perform image encoding and reconstruction processing by inputting the multiple sample images into an initial differentiable encoder, thereby obtaining multiple sample reconstructed images and multiple prediction code rates.

[0278] The coding loss information determination module is configured to determine coding loss information based on the reconstructed image from the plurality of samples, the plurality of sample images, and the plurality of predicted code rates.

[0279] The iterative training module of the differentiable encoder is configured to train the initial differentiable encoder using the encoding loss information until the loss condition is met, and then use the initial differentiable encoder corresponding to the loss condition as the differentiable encoder.

[0280] In one possible implementation, the resolution of the sample filtered image is lower than the resolution of the sample image frame; the apparatus may further include:

[0281] The first upsampling module is configured to perform upsampling processing on multiple reconstructed image frames of the samples to obtain multiple upsampled images;

[0282] The image distortion information determination module is further configured to determine the image distortion information based on the plurality of upsampled images and the plurality of sample image frames.

[0283] In one possible implementation, the first upsampling module may include:

[0284] The upsampling unit is configured to input multiple reconstructed image frames from the samples into a preset upsampling network for upsampling processing to obtain the multiple upsampled images.

[0285] The device may further include:

[0286] The upsampling network training module is configured to train the preset upsampling network using the loss information until the loss condition is met, and then use the preset upsampling network that meets the loss condition as the target upsampling network.

[0287] This disclosure also provides a video encoding apparatus, which may include:

[0288] The second filtering module is configured to perform filtering processing on multiple image frames to be encoded in the video to be encoded based on the target filtering network, so as to obtain at least one target filtered image corresponding to each of the image frames to be encoded.

[0289] The encoding module is configured to input the target filtered image into the target encoder for image encoding processing to obtain the encoded image data corresponding to each of the plurality of image frames to be encoded.

[0290] The target filtering network is obtained according to the training method described above.

[0291] In one possible implementation, the plurality of image frames to be encoded includes a target reference image frame and adjacent image frames that are temporally adjacent to the target reference image frame; the second filtering module may include:

[0292] The first filtered image acquisition unit is configured to input the target reference image frame and its temporally adjacent neighboring image frames into the target filtering network, perform filtering processing on the target reference image frame, and obtain the first filtered image corresponding to the target reference image frame.

[0293] Accordingly, the encoding module may include:

[0294] The first encoding unit is configured to input adjacent image frames that are temporally adjacent to the first filtered image and the target reference image frame into the target encoder for image encoding processing, thereby obtaining encoded image data corresponding to each of the plurality of image frames to be encoded.

[0295] In one possible implementation, the video encoding device may further include:

[0296] The initial image frame acquisition module is configured to acquire multiple initial image frames to be encoded in the video to be encoded;

[0297] The image frame acquisition module is configured to perform filtering processing on the initial image frame input filter to obtain the plurality of image frames to be encoded.

[0298] In one possible implementation, the video encoding device may further include:

[0299] The secondary filtering module is configured to perform filtering processing on the target filtered image input to the filter to obtain a secondary filtered image;

[0300] The encoding module may include:

[0301] The second encoding unit is configured to input the secondary filtered image into the target encoder for image encoding processing to obtain the encoded image data corresponding to each of the plurality of image frames to be encoded.

[0302] In one possible implementation, the resolution of the target filtered image is lower than the resolution of the image frame to be encoded; the video encoding apparatus may further include:

[0303] The second upsampling module is configured to perform upsampling by inputting the encoded image data into the target upsampling network for upsampling processing to obtain target encoded data; the target upsampling network can be the target upsampling network trained above.

[0304] Regarding the apparatus in the above embodiments, the specific manner in which each module performs its operation has been described in detail in the embodiments related to the method, and will not be elaborated upon here.

[0305] Figure 12 This is a block diagram illustrating an electronic device for video encoding according to an exemplary embodiment. The electronic device may be a terminal, and its internal structure diagram may be as follows: Figure 12As shown, the electronic device includes a processor, memory, network interface, display screen, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage medium. The network interface is used to communicate with external terminals via a network connection. When the computer program is executed by the processor, it implements a video encoding method. The display screen can be a liquid crystal display (LCD) or an e-ink display. The input devices can be a touch layer covering the display screen, buttons, a trackball, or a touchpad mounted on the device's casing, or an external keyboard, touchpad, or mouse.

[0306] Those skilled in the art will understand that Figure 12 The structure shown is merely a block diagram of a portion of the structure related to the present disclosure and does not constitute a limitation on the electronic device to which the present disclosure is applied. A specific electronic device may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0307] Figure 13 This is a block diagram illustrating an electronic device for training a filtering network according to an exemplary embodiment. The electronic device may be a server, and its internal structure diagram may be as follows: Figure 13 As shown, the electronic device includes a processor, memory, and a network interface connected via a system bus. The processor provides computational and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The network interface is used to communicate with external terminals via a network connection. When the computer program is executed by the processor, it implements a method for training a filtering network.

[0308] Those skilled in the art will understand that Figure 13 The structure shown is merely a block diagram of a portion of the structure related to the present disclosure and does not constitute a limitation on the electronic device to which the present disclosure is applied. A specific electronic device may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0309] In an exemplary embodiment, an electronic device is also provided, including: a processor; and a memory for storing processor-executable instructions; wherein the processor is configured to execute the instructions to implement a training method for a filtering network or a video coding method as described in the embodiments of this disclosure.

[0310] In an exemplary embodiment, a computer-readable storage medium is also provided, which, when executed by a processor of an electronic device, enables the electronic device to perform a training method for a filtering network or a video coding method according to embodiments of the present disclosure. The computer-readable storage medium may be a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, or optical data storage device, etc.

[0311] In an exemplary embodiment, a computer program product containing instructions is also provided, which, when run on a computer, causes the computer to perform the training method or video coding method of the filtering network in the embodiments of this disclosure.

[0312] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. This computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and RAMbus dynamic RAM (RDRAM), etc.

[0313] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the following claims.

[0314] It should be understood that this disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this disclosure is limited only by the appended claims.

Claims

1. A method of training a filter network, the method comprising: The method comprises the following steps: filtering a plurality of sample image frames related in time domain in a sample video based on an initial filtering network to obtain at least one sample filtered image corresponding to each of the sample image frames; inputting the at least one sample filtered image into a differentiable encoder to perform image encoding and reconstruction processing to obtain a sample reconstructed image frame corresponding to each of the plurality of sample image frames and a sample code rate corresponding to each of the plurality of sample image frames; the differentiable encoder refers to functions used by each module for encoding and reconstruction are differentiable or each module is a neural network; determining image distortion information according to a plurality of sample reconstructed image frames and a plurality of sample image frames; determining loss information based on the image distortion information and a plurality of sample code rates; training the initial filtering network using the loss information until a loss condition is met, and using the initial filtering network corresponding to the loss condition as a target filtering network; wherein the plurality of sample image frames comprise a sample reference image frame and adjacent image frames of the sample reference image frame; the sample filtered image comprises a first sample filtered image; the sample reconstructed image frame comprises a first sample reconstructed image frame and a second sample reconstructed image frame; the sample code rate comprises a first sample code rate and a second sample code rate; the filtering a plurality of sample image frames related in time domain in a sample video based on an initial filtering network to obtain at least one sample filtered image corresponding to each of the sample image frames comprises: inputting the sample reference image frame and the adjacent image frames of the sample reference image frame into the initial filtering network to perform filtering processing on the sample reference image frame to obtain the first sample filtered image corresponding to the sample reference image; the inputting the at least one sample filtered image into a differentiable encoder to perform image encoding and reconstruction processing to obtain a sample reconstructed image frame corresponding to each of the plurality of sample image frames and a sample code rate corresponding to each of the plurality of sample image frames comprises: inputting the first sample filtered image into the differentiable encoder to perform encoding and reconstruction processing on the first sample filtered image to obtain a first sample reconstructed image frame and a first sample code rate; inputting the first sample reconstructed image frame and the adjacent image frames into the differentiable encoder to perform encoding and reconstruction processing on the adjacent image frames to obtain a second sample reconstructed image frame corresponding to each of the adjacent image frames and a second sample code rate.

2. The method of claim 1, wherein, The differentiable encoder comprises a differentiable motion estimation network, a motion compensation network, a transform network, a quantization network, a dequantization network, an inverse transform network, a loop filtering network, and an entropy encoding network; the inputting the first sample filtered image into the differentiable encoder to perform encoding and reconstruction processing on the first sample filtered image to obtain a first sample reconstructed image frame and a first sample code rate comprises: performing motion estimation, motion compensation, transform encoding, and quantization processing on the first sample filtered image in sequence based on the motion estimation network, the motion compensation network, the transform network, and the quantization network to obtain image quantization data; Input the image quantization data and the motion vector output by the motion estimation network into the entropy encoding network for encoding processing to obtain the first sample code rate; Input the image quantization data into the inverse quantization network, the inverse transform network, the motion compensation network and the loop filtering network for image reconstruction processing to obtain the first sample reconstructed image frame.

3. The method of claim 1, wherein, The inputting the sample reference image frame and the adjacent image frame of the sample reference image frame into the initial filtering network for filtering processing on the sample reference image frame to obtain the first sample filtered image corresponding to the sample reference image frame comprises: Performing motion compensation processing on the adjacent image frame to obtain a target adjacent image frame; Inputting the sample reference image frame and the target adjacent image frame into the initial filtering network for filtering processing on the sample reference image frame to obtain the first sample filtered image corresponding to the sample reference image frame.

4. The method of claim 1, wherein, The sample filtered image comprises a second sample filtered image; the filtering processing on the multiple sample image frames related in the time domain in the sample video based on the initial filtering network to obtain the sample filtered image corresponding to each of the at least one sample image frame comprises: Inputting the multiple sample image frames into the initial filtering network for filtering processing on the multiple sample image frames to obtain the second sample filtered image corresponding to each of the multiple sample image frames; The inputting at least one sample filtered image into the derivable encoder for image encoding and reconstruction processing to obtain the sample reconstructed image frame corresponding to each of the multiple sample image frames and the sample code rate corresponding to each of the multiple sample image frames comprises: Inputting multiple second sample filtered images into the derivable encoder for encoding and reconstruction processing on the multiple second sample filtered images to obtain the multiple sample reconstructed image frames and the multiple sample code rates.

5. The method of claim 1, wherein, The filtering processing on the multiple sample image frames related in the time domain in the sample video based on the initial filtering network to obtain the sample filtered image corresponding to each of the at least one sample image frame comprises: Performing prediction processing on the respective filtering weights of the multiple sample image frames related in the time domain in the sample video based on the initial filtering network to obtain the sample filtering weight corresponding to each of the at least one sample image frame; Performing weighting processing on the at least one sample image frame according to the sample filtering weight to obtain the sample filtered image corresponding to each of the at least one sample image frame.

6. The method according to any one of claims 1 to 5, characterized in that, The determining image distortion information according to the multiple sample reconstructed image frames and the multiple sample image frames comprises: Determining image distortion sub-information between each sample image frame and the corresponding sample reconstructed image frame; Performing weighting processing on the multiple image distortion sub-informations based on the distortion weight corresponding to each of the multiple sample image frames to obtain the image distortion information.

7. The method of claim 6, wherein, The determining loss information based on the image distortion information and the multiple sample code rates comprises: Performing weighting processing on the multiple sample code rates based on the code rate weight corresponding to each of the multiple sample image frames to obtain a sample joint code rate; The image distortion information and the sample joint code rate are added to obtain the loss information.

8. The method according to any one of claims 1 to 5, characterized in that, The method further comprises: obtaining a plurality of sample images; inputting the plurality of sample images into an initial derivable encoder for image encoding and reconstruction processing to obtain a plurality of sample reconstructed images and a plurality of predicted code rates; determining encoding loss information according to the plurality of sample reconstructed images, the plurality of sample images, and the plurality of predicted code rates; training the initial derivable encoder using the encoding loss information until a loss condition is met, and using the initial derivable encoder corresponding to the time when the loss condition is met as the derivable encoder.

9. The method of claim 1, wherein, The resolution of the sample filtered image is lower than the resolution of the sample image frame; before determining the image distortion information according to the plurality of sample reconstructed image frames and the plurality of sample image frames, the method further comprises: upsampling the plurality of sample reconstructed image frames to obtain a plurality of upsampled images; determining the image distortion information according to the plurality of upsampled images and the plurality of sample image frames. The upsampling of the plurality of sample reconstructed image frames to obtain a plurality of upsampled images comprises:

10. The method of claim 9, wherein, inputting the plurality of sample reconstructed image frames into a preset upsampling network for upsampling processing to obtain the plurality of upsampled images; The method further comprises training the preset upsampling network using the loss information until a loss condition is met, and using the preset upsampling network corresponding to the time when the loss condition is met as a target upsampling network. comprises:

11. A method of video coding, the method comprising: filtering a plurality of to-be-encoded image frames of a to-be-encoded video based on a target filtering network to obtain a target filtered image corresponding to each of the plurality of to-be-encoded image frames; inputting the target filtered image into a target encoder for image encoding processing to obtain encoded image data corresponding to each of the plurality of to-be-encoded image frames; The target filtering network is trained according to the method of any one of claims 1-10. The plurality of to-be-encoded image frames comprise a target reference image frame and adjacent image frames adjacent to the target reference image frame in the time domain; 12. The method of claim 11, wherein, The filtering of the plurality of to-be-encoded image frames of the to-be-encoded video based on the target filtering network to obtain a target filtered image corresponding to each of the plurality of to-be-encoded image frames comprises: inputting the target reference image frame and the adjacent image frames adjacent to the target reference image frame in the time domain into the target filtering network to filter the target reference image frame and obtain a first filtered image corresponding to the target reference image frame; Correspondingly, the inputting of the target filtered image into the target encoder for image encoding processing to obtain encoded image data corresponding to each of the plurality of to-be-encoded image frames comprises: inputting the first filtered image and the adjacent image frames adjacent to the target reference image frame in the time domain into the target encoder for image encoding processing to obtain encoded image data corresponding to each of the plurality of to-be-encoded image frames. ​ 13. The method according to claim 11 or 12, characterized in that, Before the step of filtering a plurality of to-be-encoded image frames of a to-be-encoded video based on a target filtering network to obtain at least one target filtered image corresponding to each of the to-be-encoded image frames, the method further comprises: obtaining a plurality of initial image frames to be encoded in the to-be-encoded video; inputting the initial image frames into a filter to perform filtering processing to obtain the plurality of to-be-encoded image frames.

14. The method of claim 11 or 12, wherein, After the step of filtering a plurality of to-be-encoded image frames of a to-be-encoded video based on a target filtering network to obtain at least one target filtered image corresponding to each of the to-be-encoded image frames, the method further comprises: inputting the target filtered image into a filter to perform filtering processing to obtain a secondary filtered image; the step of inputting the target filtered image into a target encoder to perform image encoding processing to obtain encoded image data corresponding to each of the plurality of to-be-encoded image frames, comprises: inputting the secondary filtered image into the target encoder to perform image encoding processing to obtain encoded image data corresponding to each of the plurality of to-be-encoded image frames.

15. The method of claim 11, wherein, The resolution of the target filtered image is lower than the resolution of the to-be-encoded image frames; the method further comprises: inputting the encoded image data into a target up-sampling network to perform up-sampling processing to obtain target encoded data; the target up-sampling network is obtained by training according to the method of claim 10.

16. An apparatus for training a filter network, characterized by comprises: a first filtering module configured to perform filtering processing on a plurality of sample image frames related in the time domain in a sample video based on an initial filtering network to obtain at least one sample filtered image corresponding to each of the sample image frames; an encoding reconstruction module configured to perform image encoding and reconstruction processing by inputting at least one sample filtered image into a differentiable encoder to obtain a sample reconstructed image frame corresponding to each of the plurality of sample image frames and a sample code rate corresponding to each of the plurality of sample image frames; the differentiable encoder refers to a function used by each module for encoding and reconstruction being differentiable or the each module being a neural network; an image distortion information determination module configured to determine image distortion information according to a plurality of sample reconstructed image frames and a plurality of sample image frames; a loss determination module configured to determine loss information based on the image distortion information and a plurality of sample code rates; a training module configured to train the initial filtering network using the loss information until a loss condition is met, and use the initial filtering network corresponding to the time when the loss condition is met as a target filtering network; wherein the plurality of sample image frames comprise a sample reference image frame and a neighboring image frame of the sample reference image frame; the sample filtered image comprises a first sample filtered image; the sample reconstructed image frame comprises a first sample reconstructed image frame and a second sample reconstructed image frame; the sample code rate comprises a first sample code rate and a second sample code rate; and the first filtering module comprises: The first filtering unit is configured to input the sample reference image frame and a neighboring image frame of the sample reference image frame into the initial filtering network, perform filtering processing on the sample reference image frame, and obtain a first sample filtered image corresponding to the sample reference image. The encoding reconstruction module comprises: The reference frame encoding reconstruction unit is configured to input the first sample filtered image into the differentiable encoder, perform encoding and reconstruction processing on the first sample filtered image, and obtain a first sample reconstructed image frame and a first sample code rate. The neighboring frame encoding reconstruction unit is configured to input the first sample reconstructed image frame and the neighboring image frame into the differentiable encoder, perform encoding and reconstruction processing on the neighboring image frame, and obtain a second sample reconstructed image frame and a second sample code rate corresponding to each of the neighboring image frames.

17. A video encoding apparatus, comprising: Comprise: The second filtering module is configured to perform filtering processing on a plurality of to-be-encoded image frames of a to-be-encoded video based on a target filtering network, and obtain at least one target filtered image corresponding to each of the to-be-encoded image frames. The encoding module is configured to input the target filtered image into a target encoder to perform image encoding processing, and obtain encoded image data corresponding to each of the plurality of to-be-encoded image frames. The target filtering network is the target filtering network of claim 16.

18. An electronic device, comprising: Comprise: A processor; A memory for storing instructions executable by the processor; The processor is configured to execute the instructions to implement the training method of any one of claims 1-10 or the video encoding method of any one of claims 11-15.

19. A computer-readable storage medium, characterized in that, When the instructions in the computer-readable storage medium are executed by the processor of the electronic device, the electronic device can perform the training method of any one of claims 1-10 or the video encoding method of any one of claims 11-15.

Citation Information

Patent Citations

  • Video image processing method based on depth variable dimension code rate control

    CN114866782A