Multimedia data processing method, apparatus, device, and program

By employing a neural network-based image filtering processor that incorporates related image blocks and target encoding/decoding information, the method addresses the low accuracy and poor filtering effects of current image filtering methods, resulting in improved filtering accuracy and encoding quality.

JP2025516483AActive Publication Date: 2025-05-30TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2024563485
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-09-19
Filing Date
2023-07-07
Publication Date
2025-05-30
Estimated Expiration
2043-07-07

AI Technical Summary

Technical Problem

Current image filtering methods for multimedia data processing have low accuracy, resulting in poor filtering effects and reduced encoding quality.

Method used

A multimedia data processing method that utilizes related image blocks and target encoding/decoding information to perform filtering processing on target image blocks through an image filtering processor based on a neural network.

Benefits of technology

Improves the accuracy of image filtering processing, enhances the filtering effect, and increases the encoding quality of multimedia data by fully utilizing the high-quality advantages of related image blocks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025516483000001_ABST
    Figure 2025516483000001_ABST
Patent Text Reader

Abstract

Embodiments of the present application disclose a multimedia data processing method, an apparatus, a device, a storage medium, and a program product thereof. Here, the method includes: determining a first related image block related to a target image block to be filtered in multimedia data; obtaining target encoding / decoding information related to the first related image block and the target image block; and performing filtering processing on the target image block by an image filtering processor based on a neural network according to the first related image block and the target encoding / decoding information to obtain a filtered image block corresponding to the target image block.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] (Cross - reference to related applications) This application is filed based on a Chinese patent application with the application number 202211138569.5 filed with the Chinese Patent Office on September 19, 2022, claims the priority of the Chinese patent application, and all the contents of the Chinese patent application are incorporated herein by reference.

[0002] This application relates to the field of multimedia technologies, and in particular, to a multimedia data processing method, its device, equipment, storage medium, and program product.

Background Art

[0003] In the process of processing multimedia data, an encoding device can perform operations such as encoding, transformation, and quantization on an original image to obtain an encoded image, and then perform inverse quantization, inverse transformation, and prediction compensation operations on the encoded image to obtain a reconstructed image. Compared with the original image, due to the influence of quantization, part of the information of the reconstructed image is different from that of the original image, so distortion occurs in the reconstructed image. Therefore, in order to reduce the degree of distortion of the reconstructed image, it is necessary to perform filtering processing on the reconstructed image. According to practice, because the accuracy of the current filtering processing method for images is low, the filtering effect of the images is not good.

Summary of the Invention

[0004] Embodiments of this application provide a multimedia data processing method, its device, equipment, storage medium, and program product that can improve the accuracy of image filtering processing and the filtering effect of images.

[0005] In one aspect of the embodiments of this application, a multimedia data processing method executed by a computer device is provided, and the method includes: Determining a first related image block related to a target image block that is a filtering target within multimedia data; Obtaining target encoding / decoding information related to the target image block; Based on the first related image block and the target encoding / decoding information, performing a filtering process on the target image block by an image filtering processor based on a neural network to obtain a filtered image block corresponding to the target image block.

[0006] In one aspect of an embodiment of the present application, a multimedia data processing apparatus is provided, and the apparatus includes: A determination module configured to determine a first related image block related to a target image block that is a filtering target within multimedia data; An acquisition module configured to acquire target encoding / decoding information related to the target image block; A filtering module configured to perform a filtering process on the target image block by an image filtering processor based on a neural network based on the first related image block and the target encoding / decoding information to obtain a filtered image block corresponding to the target image block.

[0007] In one aspect of the present application, a computer device is provided, including a memory configured to store a computer program and a processor configured to execute the steps in the above method by calling the computer program.

[0008] In one aspect of an embodiment of the present application, a computer-readable storage medium storing a computer program including program instructions for causing a processor to execute the steps in the above method when executed by the processor is provided.

[0009] In one aspect of an embodiment of the present application, there is provided a computer program product including a computer program / instruction for causing the processor to implement the steps of the above method when executed by the processor.

[0010] In an embodiment of the present application, by introducing multi-dimensional information such as related image blocks and target encoding / decoding information, a rich amount of information is provided by the filtering process of the target image block. At the same time, the target encoding / decoding information is used to represent the degree of influence of the first related image block on the target image block. Therefore, based on the target encoding / decoding information and the image filtering processor based on a neural network, the high-quality advantages of the related image blocks are fully utilized to perform filtering processing on the reconstructed image block. In the filtering processing process for the target image block, the related image blocks are differentially used to improve the accuracy of the filtering processing of the target image block, which helps to improve the quality of the filtering processing of the target image block. That is, the accuracy of the image filtering processing can be improved, the filtering effect of the image can be improved, and the encoding quality of the multimedia data can be improved.

Brief Description of the Drawings

[0011]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8a

Figure 8b

Figure 8c

Figure 9a

Figure 9b

Figure 9c

Figure 10

Figure 11

Figure 12a

Figure 12b

Figure 13a

Figure 13b

Figure 13c

Figure 13d

Figure 13e

Figure 13f

Figure 13g

Figure 14

Figure 15

Embodiments for Carrying Out the Invention

[0012] To more clearly explain the technical solutions of the embodiments of the present application or the prior art, the following briefly introduces the drawings used in the description of the embodiments or the prior art. The drawings described below are only some embodiments of the present application, and it is obvious to those skilled in the art that other drawings can be obtained according to these drawings without creative work.

[0013] Hereinafter, with reference to the drawings of the embodiments of the present application, the technical solutions of the embodiments of the present application will be clearly and completely described. Obviously, the described embodiments are only a part of the embodiments of the present application, not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of the present application without creative efforts shall be included in the protection scope of the present invention.

[0014] Embodiments of the present application relate to a technology for processing multimedia data. Here, multimedia data (also referred to as media data) refers to composite data formed by media data such as text, graphics, images, audio, video, dynamic images, etc. whose contents are related to each other. The multimedia data referred to in the embodiments of the present application mainly includes image data consisting of images, or video data consisting of images and audio, etc. The processing process for the multimedia data according to the embodiments of the present application mainly includes media data collection, media data encoding, media data file packaging, media data file transmission, media data decoding, and final data display. When the multimedia data is video data, the complete processing process of the video data can be as shown in FIG. 1, specifically, it may include video collection, video encoding, video file packaging, video transmission, video file unpackaging, video decoding, and final video display.

[0015] Video collection is used to convert analog video into digital video and save it in the form of digital video files. That is, video collection can convert video signals into binary digital information. Here, the binary information converted from the video signal is a binary data stream, and this binary information is also called the code stream or bitstream of the video signal. Video encoding is to use compression technology to convert a file in the original video format into a file in another video format. The generation of video media content mentioned in the embodiments of this application includes the actual scene generated by camera collection and the scene of screen content generated by a computer. From the perspective of the acquisition method of video signals, video signals can be divided into two methods: those captured by a camera and those generated by a computer. Due to differences in statistical characteristics, the corresponding compression encoding methods may also be different. Taking the modern mainstream video encoding technologies, such as the international video encoding standard (HEVC: High Efficiency Video Coding) / H.265, the international video encoding standard (VVC: Versatile Video Voding) / H.266, and the Chinese national video encoding standard (AVS: Audio Video Coding Standard), or AVS3 (the third-generation video encoding standard introduced by the AVS standard group) as examples, a hybrid encoding framework is adopted to perform the following series of operations and processes on the input original video signal. Specifically, it is as shown in Figure 2. (i) Block partition structure: The input multimedia data frame (for example, one video frame in the video data) is divided into several non-overlapping processing units based on the size of one processing unit, and the same compression operation is performed on each processing unit. In one embodiment, this processing unit is called a coding tree unit (CTU) or the largest coding unit (LCU).Here, if the CTU is further subdivided, finer division can be performed to obtain one or more basic coding units called coding units (CUs). Each CU is the most fundamental element in one coding stage. In another example, this processing unit is also called a coding tile (a rectangular region of a multimedia data frame that can be decoded and encoded independently). Here, if the coding tile is further subdivided, finer division can be performed to obtain one or more superblocks (SBs) (which are the starting points of block division and can continue to be divided into multiple subblocks), and then the superblock is further divided to obtain one or more image blocks (Bs). Each image block is the most fundamental element in one coding stage.

[0016] Here, the relationship between the LCU (or CTU) and the CU is as shown in FIG. 3. As can be seen from FIG. 3, one image frame in the multimedia data includes one or more maximum coding units, and one maximum coding unit includes at least two coding units. Understandably, in the coding process for multimedia data, if the coding device performs block splitting processing on the image frame in the multimedia data, one image block in the embodiments of the present application may be referred to as the maximum coding unit (LCU) or the coding unit (CU) in one frame of the image in the multimedia data. In the coding process for multimedia data, if the coding device does not perform block splitting processing on the image frame in the multimedia data, one image block in the embodiments of the present application may be referred to as one frame of the image in the multimedia data. (ii) Predictive Coding: It includes methods such as intra prediction and inter prediction. The original video signal can obtain a residual video signal by predicting the selected coded video signal. On the coding side, for the current coded image block (i.e., the image block to be decoded), it is necessary to select the most suitable one from various possible predictive coding modes and notify the decoding side.

[0017] a. Intra (picture) Prediction: The predicted signal is from the coded and reconstructed regions within the same image.

[0018] b. Inter (picture) Prediction: The predicted signal is from another coded image (referred to as the reference image) different from the current image.

[0019] (iii) Transform & Quantization: The residual video signal is transformed into the transform domain by a transform operation such as the Discrete Fourier Transform (DFT) or the Discrete Cosine Transform (DCT), and these are called transform coefficients. DCT is a signal in the transform domain of a subset of the DFT, and further performs a lossy quantization operation so that a certain amount of information is lost, making the quantized signal advantageous for compressed representation.

[0020] In some video coding standards, there may be multiple selectable transform methods. Therefore, on the encoding side, it is necessary to select one of them for the currently encoded image block and notify the decoding side. The fineness of quantization is usually determined by the Quantization Parameter (QP). When the value of QP is large, it means that coefficients with larger values are quantized to the same output. Therefore, the distortion increases and the code rate decreases. Conversely, when the value of QP is small, it means that coefficients with smaller values are quantized to the same output. Therefore, while the distortion decreases, the code rate increases.

[0021] (iv) Entropy Coding or Statistical Coding: The quantized transform domain signal is statistically compression-encoded based on the frequency of occurrence of each value, and finally, a binary (0 or 1) compressed code stream is output. At the same time, other information is generated by encoding, such as the selected mode, motion vector, etc. In order to reduce the code rate, it is also necessary to perform entropy coding on them.

[0022] Statistical coding is a type of reversible coding method that can effectively reduce the code rate required to represent the same signal. Common statistical coding methods include variable length coding (VLC) or content adaptive binary arithmetic coding (CABAC).

[0023] (v) Loop Filtering: The encoded image (i.e., multimedia data frame) can obtain a reconstructed decoded image through operations of inverse quantization, inverse transformation, and prediction compensation (reverse operations of (ii) to (iv) above). Compared with the original image, due to the influence of quantization, part of the information in the reconstructed image is different from the original image, resulting in distortion. By performing a filtering operation on the reconstructed image, for example, using filters such as deblocking, sample adaptive offset (SAO), or adaptive loop filter (ALF), the degree of distortion caused by quantization can be effectively reduced. Since these filtered reconstructed images are used as references for subsequent encoded images to predict subsequent signals, the above filtering operation is also called loop filtering and the filtering operation within the coding loop.

[0024] Figure 2 shows the basic process of a video encoder. In Figure 2, the k-th CU (denoted as S k [x, y]) is taken as an example for explanation. Here, k is a positive integer greater than or equal to 1 and less than or equal to the number of CUs in the input current image, and S k [x, y] represents the pixel point at coordinates [x, y] in the k-th CU. x represents the horizontal coordinate of the pixel point, y represents the vertical coordinate of the pixel point, and S k [x, y] obtains a prediction signal S^ k [x, y] through one of the preferred processes such as motion compensation or intra prediction, and S k[x, y] is subtracted from S^ k [x, y] to obtain a residual signal U k [x, y]. Then, the residual signal U k [x, y] is subjected to transformation and quantization. The data output by quantization is output to two different processes. One is sent to an entropy encoder for entropy encoding, and the encoded code stream is output to a buffer and stored, waiting to be output. The other performs inverse quantization and inverse transformation to obtain a signal U’ k [x, y]. The signal U’ k [x, y] is added to S^ k [x, y] to obtain a new prediction signal S* k [x, y]. S* k [x, y] is sent to and stored in the buffer of the current image. S* k [x, y] is used to obtain f(S* k [x, y]) by intra-image prediction. S* k [x, y] is loop-filtered to obtain S’ k [x, y]. To generate a reconstructed video, S’ k [x, y] is sent to and stored in the decoded image buffer. S’ k [x, y] is obtained by motion-compensated prediction of S’ r [x + m x,y + m y . S’ r [x + m x,y + m y indicates a reference block, and m x and m y respectively indicate the horizontal and vertical components of the motion vector.

[0025] After encoding multimedia data, it is necessary to package the encoded data stream and transmit it to the user. Video file packaging refers to storing compressed and encoded video and audio in a single file according to a certain format according to a packaging format (or container, or file container). Common packaging formats include the Audio Video Interleaved format (AVI) or the media file format based on the International Standard Organization (ISO) standard (ISOBMFF: ISO Based Media File Format). Here, ISOBMFF is the packaging standard for media files, and the most typical ISOBMFF file is the Moving Picture Experts Group 4 (MP4) file.

[0026] The packaged file is transmitted to a decoding device (i.e., a user terminal) via video. After the decoding device performs reverse operations such as unpacking and decoding, the final video content can be presented on the decoding device. Here, the packaged file can be sent to the decoding device via a transmission protocol, and the transmission protocol may be, for example, Dynamic Adaptive Streaming over HTTP (DASH) based on HTTP. It is an adaptive bitrate streaming technology. By using DASH for transmission, high-quality streaming media can be distributed over the Internet via a conventional HTTP network server. In DASH, media fragment information is described using Media Presentation Description (MPD) within DASH. Also, in DASH, a combination of one or more media components, such as a video file of a certain resolution, can be regarded as one representation. The multiple representations included can be regarded as a set of one video stream (Adaptation Set), and one DASH may include one or more Adaptation Sets.

[0027] It is understandable that the process of unpacking the file packaging of the decoding device is the reverse of the above file packaging process. The decoding device unpacks the packaging file according to the file format requirements during packaging to obtain the audio code stream and the video code stream. The decoding process of the decoding device is also the reverse of the encoding process, and the decoding device can decode the audio code stream to restore the audio content. As can be seen from the above encoding process, on the decoding side, for each CU, after the decoder obtains the compressed code stream, first, entropy decoding is performed to obtain various mode information and quantized transform coefficients. From each coefficient, an inverse quantization and an inverse transform are used to obtain the residual signal. On the other hand, based on the known encoding mode information, the prediction signal corresponding to the CU can be obtained, and the two can be added to obtain the reconstructed signal. Finally, the reconstructed value of the decoded image requires a loop filtering operation to generate the final output signal.

[0028] As will be understood hereinafter, in both the encoding side and the decoding side described above, it is necessary to perform filtering processing on the reconstructed image. Since the accuracy of the current filtering processing method for images is low, the filtering effect of the image is not good. In view of this, in the embodiments of the present application, by introducing multi-dimensional information such as related image blocks and target encoding / decoding information, while providing a rich amount of information through the filtering process of the target image block, the target encoding / decoding information is used to represent the degree of influence of the first related image block on the target image block. Therefore, by using the target encoding / decoding information, the high-quality advantages of the related image blocks are fully utilized to perform filtering processing on the reconstructed image block. In the filtering processing process for the target image block, the related image blocks are differentially used to improve the accuracy of the filtering processing of the target image block and contribute to improving the quality of the filtering processing of the target image block. That is, the accuracy of the filtering processing of the image can be improved, the filtering effect of the image can be improved, and furthermore, the encoding quality of the multimedia data can be improved.

[0029] Furthermore, referring to FIG. 4, it is an exemplary flowchart of a multimedia data processing method according to an embodiment of the present application. The method can be executed by a computer device, and the computer device may refer to an encoding device or a decoding device. As shown in FIG. 4, the method may at least include the following S101 to S103.

[0030] In S101, a first related image block related to a target image block to be filtered in the multimedia data is determined.

[0031] In S102, target encoding / decoding information related to the target image block is obtained.

[0032] In step S101 and step S102, the computer device can determine a reconstructed and unfiltered image block in the multimedia data as a target image block to be filtered, or can determine a preliminarily filtered reconstructed image block in the multimedia data as a target image block to be filtered. Here, the preliminary filtering process may refer to a conventional filtering process method. Further, a first related image block related to the target image block can be obtained, and target encoding / decoding information related to the first related image block and the target image block can be obtained. The target encoding / decoding information includes encoding / decoding information corresponding to the first related image block, and the target encoding / decoding information may further include encoding / decoding information corresponding to the target image block. The first related image block belongs to a related image having an encoding reference relationship with the target image block, that is, the first related image block belongs to a reference image having an encoding reference relationship with the target image block.

[0033] It is understandable that multimedia data includes images of multiple frames. The images of multiple frames are also called a sequence of multimedia data. An image of one frame includes one or more slices. One slice includes one or more image blocks. When an image of one frame includes one slice and one slice includes one image block, the image block refers to the image corresponding to the image block. According to the encoding method of slices, the slices of an image include complete intra-coded slices and incomplete intra-coded slices. A complete intra-coded slice means that the encoding method of all image blocks within the slice is the intra-coding (i.e., intra-prediction) method. An incomplete intra-coded slice means that the encoding method of the image blocks within the slice may be the inter-coding (i.e., inter-prediction) method. When all slices within an image belong to complete intra-coded slices, the image is also called an I-frame or a key frame. When there are slices in the image that belong to incomplete intra-coded slices, the image may be a B-frame or a P-frame. The uncoded image blocks within an image are also called original image blocks. The coded original image blocks within an image are also called coded image blocks (or predicted image blocks). The reconstructed coded image blocks within an image are also called reconstructed image blocks. The unreconstructed reconstructed image blocks or the initially filtered reconstructed image blocks within an image are also called target image blocks to be filtered. The filtered target image blocks of the present solution within an image are also called filtered image blocks.

[0034] As can be understood, in multimedia data encoding technology, the time-domain layer division technology is also involved. According to the dependency relationship during encoding, different image frames can be divided into different time-domain layers. Specifically, using the time-domain hierarchical technology, the images in the multimedia data are divided into time-domain layers. When an image frame is divided into a lower layer, it is not necessary to refer to the image frames of the higher layer during encoding. Here, the time-domain layer of an image can be used to represent the encoding order of the image. For example, the lower the time-domain layer, the earlier the encoding order of the image, and the higher the time-domain layer, the later the encoding order of the image. The time-domain layer of an image can also be used to represent, for example, the image distance between the image and its reference image. The larger the time-domain layer of the image, the closer the image distance between the image and its reference image, and the smaller the time-domain layer of the image, the farther the image distance between the image and its reference image. The time-domain layers of all the image blocks within the same image are the same, and the time-domain layer of an image block refers to the time-domain layer of the image to which the image block belongs.

[0035] It is understandable that the target image block and the first related image block may belong to the same image in the multimedia data or may belong to different images in the multimedia data. For example, the image to which the target data block belongs is an I-frame, and the I-frame refers to a key frame. Such an image frame of this media type only allows the use of the intra-coding method without the need to be coded depending on other image frames. That is, when the coding method of the target image block is the intra-coding method, the target image block itself can be used as the first related image block. As shown in FIG. 5, the multimedia data includes 9 frames of images, and the coding methods of these 9 frames of images are all intra-coding methods. The coding processes of the image blocks in each frame of the image all depend on the image blocks in their respective images. The numbers above each frame of the image are for indicating the coding order of each frame of the image. There is no coding reference relationship between the images of each frame, and the temporal layer IDs of the images of each frame are all 0.

[0036] In another example, the image to which the target data block belongs is a B frame or a P frame, also called an inter-coded frame. Such an image frame allows the use of both inter-coding and intra-coding methods. In this case, the coding method of the target image block is a non-intra coding method. In this case, the first related image block belongs to a related image that has a coding reference relationship with the target image block. As shown in FIG. 6, the multimedia data includes 9 frames of images. Here, the media type of 8 frames of images is B frame, and the media type of 1 frame of image is I frame. The numbers above the images of each frame are for indicating the coding order of the images of each frame, and the arrows between the images of each frame are used to represent the coding reference relationship between the images of each frame. The temporal layer ID of each frame of image in FIG. 6 is all 0. The media type of the image located at the head of the coding order is I frame, and the media type of the image located behind the head of the coding order is B frame. The coding process of each B frame all refers to the first I frame and the corresponding adjacent image frame (i.e., the image frame whose coding order is located before it). That is, when the target image block belongs to the B frame in FIG. 6, the first related image block belongs to the first I frame or the adjacent image of the image to which the target image block belongs. When the target image block belongs to the image located at the third position in the coding order, the number of the first related image blocks related to the target image block is two. For example, the first related image block 1 and the first related image block 2. The first related image block 1 belongs to the first I frame, and the first related image block 2 belongs to the image whose coding order is located at the second position.

[0037] In another example, as shown in FIG. 7, the multimedia data includes 9 frames of images, where the media type of 8 frames of images is B-frame, the media type of 1 frame of image is I-frame, the numbers above the images of each frame are used to indicate the encoding order of the images of each frame, and the arrows between the images of each frame are used to represent the encoding reference relationship between the images of each frame. In FIG. 7, the media type of the image located at the head of the encoding order is I-frame, and the media type of the images located behind the head of the encoding order is B-frame. The temporal region layers of the image located at the head of the encoding order and the image located at the second position in the encoding order are both 0, indicating Layer0, the temporal region layer of the image located at the third position in the encoding order is 1, indicating Layer1. The temporal region layers of the image located at the fourth position in the encoding order and the image located at the seventh position in the encoding order are both 2, indicating Layer2, and the temporal region layers of the image located at the fifth position in the encoding order, the image located at the sixth position in the encoding order, the image located at the eighth position in the encoding order, and the image located at the ninth position in the encoding order are all 3, indicating Layer3. The encoding process of the image with a lower temporal region layer does not depend on the image with a higher temporal region layer, and the encoding process of the image with a higher temporal region layer can depend on the image with a lower temporal region layer.

[0038] It should be understood that the high or low of the time domain layers mentioned in the embodiments of the present application is a relative concept, like the four time domain layers of Layer0 to Layer3 determined in FIG. 7. For the time domain layer of Layer0, Layer1 to Layer3 are all high time domain layers. For the time domain layer of Layer1, the time domain layer of Layer3 is the high time domain layer of Layer1, and the time domain layer of Layer0 is the low time domain layer of Layer1. In FIG. 7, when the target image block belongs to the image located first in the encoding order, the related image blocks related to the target image block also belong to the image located first in the encoding order. When the target image block belongs to the image located after the first one (i.e., the image with the encoding order at the head) in the encoding order, the image to which the related image block related to the target image block belongs is different from the image to which the target image block belongs. When the target image block belongs to the image located last in the encoding order, the related image blocks related to the target image belong to the image located first in the encoding order.

[0039] It should be understood that the target encoding / decoding information is used to represent the influence degree of the first related image block on the target image block. The higher the influence degree of the first related image block on the target image block represented by the target encoding / decoding information, the more information amount of the first related image block is referred to in the encoding process of the target image block. Therefore, in the process of performing filtering processing on the target image block, more information amount of the first related image block can be referred to. The lower the influence degree of the first related image block on the target image block represented by the target encoding / decoding information, the less information amount of the first related image block is referred to in the encoding process of the target image block. Therefore, in the process of performing filtering processing on the target image block, less information amount of the first related image block can be referred to. The target decoding information enables the differential use of related image blocks and improves the accuracy of the filtering processing for the target image block.

[0040] Understandably, the target encoding / decoding information includes the encoding / decoding information corresponding to the first related image block, or the target encoding / decoding information includes the encoding / decoding information corresponding to the first related image block and the encoding / decoding information corresponding to the target image block. The encoding / decoding information corresponding to the target image block includes at least one of a sequence level quantization parameter, a slice level quantization parameter, a slice level encoding type, a block encoding type, filtering intensity image block information of the target image block, prediction image block information corresponding to the target image block, and block division image block information of the target image block. Here, the prediction image block information corresponding to the target image block includes all or part of the information of the prediction image block corresponding to the target image block. For example, the prediction image block information corresponding to the target image block includes at least one of chroma component information and luminance component information of the prediction image block corresponding to the target image block. The prediction image block corresponding to the target image block refers to the one obtained by performing predictive encoding on the original image block corresponding to the target image block. The filtering intensity image block information of the target image block refers to at least one of chroma component information and luminance component information of the filtering intensity image block information corresponding to the target image block. The block division image block information of the target image block includes all or part of the information of the image blocks in the image generated by block division. For example, the block division image block information of the target image block includes at least one of chroma component information and luminance component information of the image blocks in the image generated by block division. Here, the filtering intensity image block is an image block generated based on the filtering intensity discrimination result of the deblocking filter in the codec. The block division image block information is an image block generated based on the encoding block division result. A sequence refers to a series of encoded frames. Regarding a slice, in the encoding process of multimedia data, an image is divided into multiple slices, which is the same as a tile. If not divided, one frame may be regarded as one slice.

[0041] It is understandable that since the encoding / decoding information corresponding to the target image block has a certain influence on the image quality of the target image block, the encoding / decoding information corresponding to the target image block is used as auxiliary information for the filtering process of the target image block. The larger the amount of information included in the encoding / decoding information corresponding to the target image block, the more helpful it is for assisting the filtering process of the target image block and improving the accuracy of the filtering process of the target image block.

[0042] As can be understood, the encoding / decoding information corresponding to the first related image block includes at least one of the influencing factor of the first related image block with respect to the target image block, the reference direction of the first related image block with respect to the target image block, the block division image block information of the first related image block, the quantization parameter corresponding to the first related image block, the reconstructed image block information corresponding to the first related image block, the filtering intensity image block information corresponding to the first related image block, and the predicted image block information corresponding to the first related image block. The encoding / decoding information of the first related image block is used to represent the influence degree of the first related image block on the target image block. For example, the influencing factor is determined based on the image distance between the first related image block and the target image block, and the influencing factor is used to represent the influence degree of the first related image block on the target image block. The greater the influencing factor, the higher the influence degree of the first related image block on the target image block, and the smaller the influencing factor, the lower the influence degree of the first related image block on the target image block. The influencing factor is determined based on the image distance between the first related image block and the target image block, and there is an inverse correlation between the influencing factor and the image distance, that is, the farther the image distance between the first related image block and the target image block, the smaller the influencing factor of the first related image block on the target image block, that is, the closer the image distance between the first related image block and the target image block, the greater the influencing factor of the first related image block on the target image block.

[0043] In another example, the quantization parameter corresponding to the first related image block is used to represent the degree of influence of the first related image block on the target image block. The larger the quantization parameter corresponding to the first related image block, the less information the first related image block itself has, and the lower the degree of influence of the first related image block on the target image block. The smaller the quantization parameter corresponding to the first related image block, the more information the first related image block itself has, and the higher the degree of influence of the first related image block on the target image block. In another example, the image block information corresponding to the first related image block is used to represent the degree of influence of the first related image block on the target image block. The more information the image block information corresponding to the first related image block has, the higher the degree of influence of the first related image block on the target image block. The less information the image block information corresponding to the first related image block has, the lower the degree of influence of the first related image block on the target image block. The image block information corresponding to the first related image block includes at least one of the reconstructed image block information, block division image block information, filtering intensity image block information, and predicted image block information corresponding to the first related image block. It can be understood that the reference direction of the first related image block with respect to the target image block can be used as one piece of auxiliary information for the filtering process of the target image block, providing more auxiliary information for the filtering process of the target image block and improving the filtering effect.Understandably, when the slice where the target image block is located is an incomplete intra-coded slice, the computer device can determine the image distance between the first related image block and the target image block based on the picture order count corresponding to the first related image block and the picture order count corresponding to the target image block. The picture order count (POC) is used to represent the display order of the image to which the image block belongs. When the slice where the target image block is located is a complete intra-coded slice, a predetermined distance value is determined as the image distance between the first related image block and the target image block, and the predetermined distance value may refer to 0, 1025, or other specific values. Here, when the slice where the target image block is located is a complete intra-coded slice, the slice where the target image block is located may be referred to as an I slice. When the slice where the target image block is located is an incomplete intra-coded slice, the slice where the target image block is located may be referred to as a P slice or a B slice. By generating the image distance between the target image block and the related image block using the above method, it is possible for the I slice, B slice, and P slice to share the above neural network-based image filtering processor for filtering processing, without the need to train a neural network-based image filtering processor for each of the I slice, B slice, and P slice, saving resources and improving the versatility of the neural network-based image filtering processor.

[0044] It is understandable that the image sequence count corresponding to the target image block refers to the image sequence count of the image frame to which the target image block belongs, and the image sequence count corresponding to the first related image block refers to the image sequence count of the image frame to which the first related image block belongs. It is understandable that the computer device can determine the image distance between the first related image block and the target image block based on the image sequence count corresponding to the first related image block and the image sequence count corresponding to the target image block through any one of the following four methods.

[0045] In Method 1, the computer device can calculate the count difference between the image sequence count corresponding to the first related image block and the image sequence count corresponding to the target image block, and determine the absolute value of the count difference as the image distance between the first related image block and the target image block.

[0046] In Method 2, the computer device obtains the count difference between the image sequence count corresponding to the first related image block and the image sequence count corresponding to the target image block, calculates the count mapping value corresponding to the absolute value of the count difference, and can determine the image distance between the first related image block and the target image block based on the count mapping value. Here, the count mapping may be (N - image distance) / M, M may be a non-zero number, N may be any value, for example, N is 1025 and M is 1024.

[0047] In Mode 3, the computer device can obtain the time-domain layer of the target image block and determine the image distance between the first related image block and the target image block based on the time-domain layer of the target image block. That is, the image distance between the first related image block and the target image block has an inverse correlation with the time-domain layer of the target image block. The higher the time-domain layer of the target image block, the closer the image distance between the first related image block and the target image block. Conversely, the lower the time-domain layer of the target image block, the farther the image distance between the first related image block and the target image block. For example, the difference between N and the time-domain layer of the target image block can be determined as the image distance between the first related image block and the target image block, and N - the time-domain layer of the target image block can be determined as the image distance between the first related image block and the target image block.

[0048] In Mode 4, the computer device can determine the level mapping value corresponding to the time-domain layer of the target image block and determine the image distance between the first related image block and the target image block based on the level mapping value. For example, the ratio of the time-domain layer of the target image block to a predetermined time-domain layer value can be determined as the level mapping value corresponding to the time-domain layer of the target image block. For example, the predetermined time-domain layer value is K, which is a non-zero number. Or, the difference between the predetermined time-domain layer value and the time-domain layer of the target image block can be determined as the level mapping value corresponding to the time-domain layer of the target image block.

[0049] It is understandable that the image block information corresponding to the above first related image block includes at least one of first color component information, second color component information, and third color component information. Here, the first color component information is used to represent the luminance of the image block, and the second color component information and the third color component information are used to represent the chroma of the image block. The image block information corresponding to the first related image block includes any one of the reconstructed image block information, block split image block information, filtering strength image block information, and predicted image block information corresponding to the first related image block.

[0050] It is understandable that when the slice where the target image block is located is an incomplete intra-coded slice, the computer device determines the reference direction of the first related image block with respect to the target image block based on the relationship between the image order count corresponding to the first related image block and the image order count corresponding to the target image block. When the slice where the target image block is located is a complete intra-coded slice, a predetermined direction value is adopted to indicate the reference direction of the first related image block with respect to the target image block. That is, by determining the reference direction of the first related image block with respect to the target image block using the above method, it is possible for I slices, B slices, and P slices to share the above neural network-based image filtering processor for filtering processing, without the need to train neural network-based image filtering processors for each of I slices, B slices, and P slices, saving resources and improving the versatility of the neural network-based image filtering processor.

[0051] For example, when the slice where the target image block is located is an incomplete intra-coded slice, and the picture order count corresponding to the first related image block is greater than the picture order count corresponding to the target image block, a first direction value is adopted to indicate the reference direction of the first related image block with respect to the target image block. When the slice where the target image block is located is an incomplete intra-coded slice, and the picture order count corresponding to the first related image block is less than the picture order count corresponding to the target image block, a second direction value is adopted to indicate the reference direction of the first related image block with respect to the target image block. The first direction value may point to a value greater than zero, and the second direction value may point to a value less than zero. The first direction value is different from the second direction value. For example, the first direction value may be 1, and the second direction value may be 0. When the slice where the target image block is located is a complete intra-coded slice, a predetermined direction value is adopted to indicate the reference direction of the first related image block with respect to the target image block. The predetermined direction value may be 0, 1, 2, or -1, etc. The predetermined direction value may be the same as the first direction value or the second direction value, or may be different from either the first direction value or the second direction value.

[0052] It can be understood that the selection method of the first related image block of the target image block may refer to selecting one or more image blocks as the first related image block from the specified reference direction. When the number of the first related image blocks is insufficient, it is filled with reusable first related image blocks. For example, the selection method of the first related image block includes three methods such as 1) a method of selecting one or two in the L0 direction, 2) a method of selecting one or two in the L1 direction, and 3) a method of selecting one or two in the L0 direction and one or two in the L1 direction. The first related image blocks include the image blocks in the first reference frame of the reference frame list in the L0 direction and the image blocks in the first reference frame of the reference frame list in the L1 direction.

[0053] Here, L0 may point in the direction where the picture order count (POC) of the picture frame to which the first related picture block belongs is smaller than the picture order count (POC) of the picture frame to which the target picture block belongs, and L1 may point in the direction where the picture order count (POC) of the picture frame to which the first related picture block belongs is larger than the picture order count (POC) of the picture frame to which the target picture block belongs. Or, both L0 and L1 point in the direction where the picture order count (POC) of the picture frame to which the first related picture block belongs is smaller than the picture order count (POC) of the picture frame to which the target picture block belongs.

[0054] It can be understood that when there is the same spatial positional relationship between the target picture block and the first related picture block, the size of the target picture block is the same as the size of the first related picture block, or the size of the target picture block is different from the size of the first related picture block. For example, the picture to which the target picture block belongs is the target picture, the picture to which the first related picture block belongs is the reference picture. All the gray areas in the target pictures in FIGS. 8a, 8b, and 8c are target picture blocks, the gray area in the reference picture in FIG. 8a is the first related picture block, and all the gray areas and hatched areas in the reference pictures in FIGS. 8b and 8c are the first related picture blocks. In FIGS. 8a, 8b, and 8c, there is the same spatial positional relationship between the target picture block and the first related picture block. Here, the size of the target picture block in FIG. 8a is the same as the size of the first related picture block, and the sizes of the target picture blocks in FIGS. 8b and 8c are smaller than the size of the first related picture block.

[0055] It is understandable that there is a different spatial positional relationship between the target image block and the first related image block. When the first related image block is determined based on the motion vector of the target image block, the size of the target image block is the same as the size of the first related image block, or the size of the target image block is different from the size of the first related image block. The motion vector of the target image block refers to one determined by the motion vectors of each encoded block within the target image block such as the average motion vector.

[0056] For example, the image to which the target image block belongs is the target image, the image to which the first related image block belongs is the reference image, the direction indicated by the arrow in FIGS. 9a, 9b, and 9c is the direction of the motion vector of the target image block, and the first related image block is obtained by selection from the reference image based on the position indicated by the motion vector of the target image block. The gray areas in the target images in FIGS. 9a, 9b, and 9c are all target image blocks, the gray area in the reference image of FIG. 9a is the first related image block, and the gray areas and hatched areas in the reference images of FIGS. 9b and 9c are all first related image blocks. In FIGS. 9a, 9b, and 9c, there is a different spatial positional relationship between the target image block and the first related image block, and the first related image block is at the position indicated by the motion vector of the target image block in the reference image. Here, the size of the target image block in FIG. 9a is the same as the size of the first related image block, and the sizes of the target image blocks in FIGS. 9b and 9c are smaller than the size of the first related image block.

[0057] In S103, based on the first related image block and the target encoding / decoding information, a filtering process is performed on the target image block by an image filtering processor based on a neural network to obtain a filtering image block corresponding to the target image block.

[0058] In an embodiment of the present application, based on the first related image block and the target encoding / decoding information, a computer device can perform filtering processing on the target image block by means of an image filtering processor based on a neural network to obtain a filtering image block corresponding to the target image block. For example, when the degree of influence of the first related image block on the target image block represented by the target encoding / decoding information is high, in the process of performing filtering processing on the target image block, the computer device can set a larger influence factor (i.e., weight) for the first related image block, which helps to increase the importance of the first related image block in the filtering processing process of the target image block. Conversely, the lower the degree of influence of the first related image block on the target image block represented by the target encoding / decoding information, the smaller influence factor (i.e., weight) the computer device can set for the first related image block in the process of performing filtering processing on the target image block, which helps to weaken the importance of the first related image block in the filtering processing process of the target image block. By means of the target encoding / decoding information and the image filtering processor based on a neural network, the high-quality advantages of the related image blocks are fully utilized to perform filtering processing on the reconstructed image blocks, and in the filtering processing process for the target image blocks, the related image blocks are differentially used to help improve the accuracy of the filtering processing of the target image blocks.

[0059] As can be understood, the computer device can perform filtering processing on the target image block based on the first related image block and the target encoding / decoding information by the above-mentioned neural network-based image filtering processor, and obtain a filtering image block corresponding to the target image block. The neural network-based image filtering processor here may refer to an in-loop filter for performing filtering processing on an image block. For example, the neural network-based image filtering processor may be a neural network-based in-loop filter (NNLF: Neural Network in-Loop Filter) or the like. As shown in FIG. 10, the computer device can input a target image block to be filtered, a first related image block, and target encoding / decoding information into the neural network-based image filtering processor. The neural network-based image filtering processor performs filtering processing on the target image block based on the first related image block and the target encoding / decoding information, and outputs a filtered image block by the neural network-based image filtering processor. The image block after the filtering processing is a filtering image block corresponding to the target image block.

[0060] As can be understood, as shown in FIG. 11, the generation process of the image filtering processor based on the neural network includes the following three steps. 1) Use the encoding software to encode the training sequence to generate a training set. 2) Based on the training set, train the initial image filtering processor based on the neural network to obtain the image filtering processor based on the neural network. 3) Integrate the image filtering processor based on the neural network into the software, and perform filtering processing on the target image block by the software in which the image filtering processor based on the neural network is integrated. In the training process of the initial image filtering processor based on the neural network, a loss function is used. The loss function measures the difference between the predicted value (the filtered image corresponding to the sample image block) and the true value (the original image block corresponding to the sample image block). The greater the loss value loss, the greater the difference. The purpose of training is to reduce loss. In the case of the encoding tool based on deep learning, the commonly used loss functions are the L1 norm loss function, the L2 norm loss function, and the smooth L1 loss function.

[0061] As can be understood, the image filtering processor based on the neural network can include a plurality of convolutional layers, a plurality of activation layers, a plurality of residual (Resblock) layers, and one data shuffle layer. Here, the convolutional layer is used to extract different input features. The activation layer performs a non-linear mapping on the output result of the convolutional layer. The Resblock layer is the basic module of the image filtering processor based on the neural network. Shuffle is used to obtain data that matches the information dimension of the original image block corresponding to the target image block through adjustment.

[0062] It can be understood that the number of channels in the convolutional layer (convolutional module) is exactly the same, that is, the number of input and output channels of all convolutional layers in the Resblock layer is the same. The Resblock layer in Figure 12a includes two 3x3 convolutional layers, and the number of input and output channels of the two 3x3 convolutional layers is the same. For example, the number of input and output channels is 64 and 64 respectively.

[0063] It can be understood that the number of channels in the convolutional layer is different, that is, the number of input and output channels of all convolutional layers in the Resblock layer may be different. The Resblock layer in Figure 12b includes two 1x1 convolutional layers and one 3x3 convolutional layer, and the number of input and output channels of the two 1x1 convolutional layers and the 3x3 convolutional layer is different. For example, the number of channels can be increased before the activation layer and decreased after the activation layer. The number of input and output channels of the first 1x1 convolutional layer is 64 and 160 respectively, the number of input and output channels of the second 1x1 convolutional layer is 160 and 64 respectively, and the number of input and output channels of the 3x3 convolution is 64 and 64 respectively.

[0064] It can be understood that the computer device can realize the filtering process for the target image block by any one of the following two methods.

[0065] In Method 1, the target encoding / decoding information includes the encoding / decoding information corresponding to the first related image block. The encoding / decoding information corresponding to the first related image block refers to the parameters correlated with the encoding / decoding process of the first related image block. The computer device can fuse the first related image block with the encoding / decoding information corresponding to the first related image block through the information fusion layer of the image filtering processor based on the neural network to obtain the target fusion data corresponding to the first related image block. Further, the computer device can perform filtering processing on the target image block through the filtering processing layer of the image filtering processor based on the neural network, based on the target fusion data corresponding to the first related image block, to obtain the filtered image block corresponding to the target image block. The encoding / decoding information corresponding to the first related image block helps to fully utilize the high-quality advantages of the first related image block and improve the accuracy and quality of the filtering processing for the target image block.

[0066] In Mode 2, the target encoding / decoding information includes the encoding / decoding information corresponding to the first related image block and the encoding / decoding information corresponding to the target image block. Specifically, the computer device can obtain the target fusion data corresponding to the first related image block by fusing the first related image block with the encoding / decoding information corresponding to the first related image block through the information fusion layer of the image filtering processor based on the neural network. Further, the computer device can perform filtering processing on the target image block by the image filtering processor based on the neural network according to the target fusion data corresponding to the first related image block and the encoding / decoding information corresponding to the target image block through the filtering processing layer of the image filtering processor based on the neural network, so as to obtain the filtered image block corresponding to the target image block. The encoding / decoding information corresponding to the first related image block and the encoding / decoding information corresponding to the target image block can make full use of the advantages of high quality of the first related image block, provide more information volume in the filtering processing process of the target image block, and help improve the accuracy and quality of the filtering processing of the target image block.

[0067] It can be understood that the information fusion layer of the image filtering processor based on the neural network may refer to the convolutional layer of the image filtering processor based on the neural network, or the information fusion layer of the image filtering processor based on the neural network may refer to the dot product module and the convolutional layer of the image filtering processor based on the neural network.

[0068] It is understandable that the number of the above-mentioned first related image blocks is M (M is an integer greater than or equal to 1), and the computer device can obtain target fusion data corresponding to the first related image block by fusing the first related image block with encoding / decoding information corresponding to the first related image block by any one of the following three fusion methods.

[0069] In fusion method 1, the image filtering processor based on a neural network includes M information fusion layers. One information fusion layer corresponds to one associated image block. When the number of the first associated image blocks is 1, the image filtering processor based on the neural network includes 1 information fusion layer. The computer device performs a convolution operation process on the first associated image block and the encoding / decoding information corresponding to the first associated image block through the information fusion layer of the image filtering processor based on the neural network, and can obtain the target fusion data corresponding to the first associated image block. When the number of the first associated image blocks is at least 2, the computer device performs a convolution operation process on the first associated image block 1 and the encoding / decoding information corresponding to the first associated image block 1 through the information fusion layer 1 of the image filtering processor based on the neural network, and can obtain the fusion data corresponding to the first associated image block 1. The information fusion layer 1 of the image filtering processor based on the neural network is the information fusion layer corresponding to the first associated image block 1 among the M information fusion layers. The computer device performs a convolution operation process on the first associated image block 2 and the encoding / decoding information corresponding to the first associated image block 2 through the information fusion layer 2 of the image filtering processor based on the neural network, and obtains the fusion data corresponding to the first associated image block 2. The information fusion layer 2 of the image filtering processor based on the neural network is the information fusion layer corresponding to the first associated image block 2 among the M information fusion layers. The above steps are repeated until the fusion data corresponding to each of the M first associated image blocks is obtained. When the fusion data corresponding to each of the M first associated image blocks is obtained, the fusion data corresponding to each of the M first associated image blocks is determined as the target fusion data corresponding to the first associated image block, and by means of the convolution operation process method, the encoding / decoding information of each first associated image block is fused with the first associated image block, which helps to improve the accuracy of the filtering process for the target image block and the quality of the filtering process.

[0070] For example, as shown in FIG. 13a, an image filtering processor based on a neural network includes at least two convolutional layers, N Resblock layers, a plurality of activation layers, and one shuffle layer. Here, the at least two convolutional layers are respectively denoted as convolutional layer 0, convolutional layer 1, convolutional layer 2, convolutional layer 3, ……, convolutional layer N, and the activation layers include activation layer 0, activation layer 1, activation layer 2, ……. When the target encoding / decoding information includes the encoding / decoding information corresponding to the first related image block, the first related image block includes related image block 0, and the encoding / decoding information corresponding to the related image block 0 is encoding / decoding information 0. The computer device inputs the related image block 0 and the encoding / decoding information 0 into the convolutional layer 0, and the convolutional layer 0 performs a convolution operation process on the related image block 0 and the encoding / decoding information 0 to obtain the fusion data corresponding to the related image block 0, and the fusion data corresponding to the related image block 0 can be determined as the target fusion data corresponding to the first related image block. The convolutional layer 1 performs a convolution operation process on the target image block, and the convolutional layer 2 performs a convolution operation on the processing result regarding the target image block output by the activation layer 1 and the processing result regarding the target fusion data output by the activation layer 0.

[0071] For example, as shown in FIG. 13b, an image filtering processor based on a neural network includes at least two convolutional layers, N Resblock layers, a plurality of activation layers, and one shuffle layer. Here, the at least two convolutional layers are respectively denoted as convolutional layer 0, convolutional layer 1, convolutional layer 2, convolutional layer 3, ……, convolutional layer N, and the activation layers include activation layer 0, activation layer 1, activation layer 2, ……. When the target encoding / decoding information includes the encoding / decoding information corresponding to the first related image block and the encoding / decoding information corresponding to the target image block, the first related image block includes related image block 0 and related image block 1. The encoding / decoding information corresponding to the related image block 0 is encoding / decoding information 0, and the encoding / decoding information corresponding to the related image block 1 is encoding / decoding information 1. Convolutional layer 0 corresponds to related image block 0, and convolutional layer 1 corresponds to related image block 1. The computer device inputs related image block 0 and encoding / decoding information 0 into convolutional layer 0, and convolutional layer 0 performs a convolution operation process on related image block 0 and encoding / decoding information 0 to obtain the fusion data corresponding to related image block 0. It inputs related image block 1 and encoding / decoding information 1 into convolutional layer 1, and convolutional layer 1 performs a convolution operation process on related image block 1 and encoding / decoding information 1 to obtain the fusion data corresponding to related image block 1. The fusion data corresponding to related image block 0 and the fusion data corresponding to related image block 1 are determined as the target fusion data corresponding to the first related image block. In FIG. 13b, the encoding / decoding information corresponding to the target image block includes a sequence level quantization parameter, a slice level quantization parameter, a slice level encoding type, a block encoding type, and a predicted image block corresponding to the target image block.

[0072] In the second fusion method, the image filtering processor based on a neural network includes one information fusion layer. The computer device performs a convolution operation on M first related image blocks and encoding / decoding information through the information fusion layer of the image filtering processor based on the neural network, and can obtain target fusion data corresponding to the first related image blocks. That is, the M first related image blocks and the encoding / decoding information are simultaneously input into the information fusion layer of the image filtering processor based on the neural network, and the information fusion layer of the image filtering processor based on the neural network performs a convolution operation on the M first related image blocks and the encoding / decoding information to obtain target fusion data corresponding to the first related image blocks. Through the convolution operation method, the encoding / decoding information of each first related image block is fused with the first related image block, which helps to improve the accuracy of the filtering process for the target image block and the quality of the filtering process.

[0073] For example, as shown in FIG. 13c, an image filtering processor based on a neural network includes at least two convolutional layers, N Resblock layers, a plurality of activation layers, and one shuffle layer. Here, the at least two convolutional layers are respectively denoted as convolutional layer 0, convolutional layer 1, convolutional layer 2, convolutional layer 3, ……, convolutional layer N, and the activation layers include activation layer 0, activation layer 1, activation layer 2, ……. When the target encoding / decoding information includes the encoding / decoding information corresponding to the first related image block and the encoding / decoding information corresponding to the target image block, the first related image block includes related image block 0 and related image block 1. The encoding / decoding information corresponding to the related image block 0 is encoding / decoding information 0, and the encoding / decoding information corresponding to the related image block 1 is encoding / decoding information 1. The computer device inputs all of related image block 0, related image block 1, encoding / decoding information 0, and encoding / decoding information 1 into convolutional layer 0, and convolutional layer 0 performs a convolutional operation process on related image block 0, related image block 1, encoding / decoding information 0, and encoding / decoding information 1, so as to obtain the target fusion data corresponding to the first related image block. In FIG. 13c, the encoding / decoding information corresponding to the target image block includes a sequence level quantization parameter, a slice level quantization parameter, a slice level encoding type, a block encoding type, and a predicted image block corresponding to the target image block.

[0074] In Fusion Method 3, the image filtering processor based on a neural network includes M information fusion layers. One information fusion layer corresponds to one associated image block. When the number of the first associated image blocks is 1, the image filtering processor based on a neural network includes 1 information fusion layer. The computer device can perform a dot product operation process on the first associated image block and the encoding / decoding information corresponding to the first associated image block through the information fusion layer of the image filtering processor based on the neural network, so as to obtain the target fusion data corresponding to the first associated image block. When the number of the first associated image blocks is at least 2, the computer device can perform a dot product operation process on the first associated image block 1 and the encoding / decoding information corresponding to the first associated image block 1 through the information fusion layer 1 of the image filtering processor based on the neural network, so as to obtain the fusion data corresponding to the first associated image block 1. The information fusion layer 1 of the image filtering processor based on the neural network is the information fusion layer corresponding to the first associated image block 1 among the M information fusion layers. The computer device can perform a dot product operation process on the first associated image block 2 and the encoding / decoding information corresponding to the first associated image block 2 through the information fusion layer 2 of the image filtering processor based on the neural network, so as to obtain the fusion data corresponding to the first associated image block 2. The information fusion layer 2 of the image filtering processor based on the neural network is the information fusion layer corresponding to the first associated image block 2 among the M information fusion layers. The above steps are repeated until the fusion data corresponding to each of the M first associated image blocks is obtained. When the fusion data corresponding to each of the M first associated image blocks is obtained, a convolution operation is performed on the fusion data corresponding to each of the M first associated image blocks to obtain the target fusion data corresponding to the first associated image block. The method of fusing the encoding / decoding information of each first associated image block with the first associated image block by the dot product operation process helps to improve the accuracy of the filtering process for the target image block and the quality of the filtering process.

[0075] For example, as shown in FIG. 13d, an image filtering processor based on a neural network includes at least two convolutional layers, N Resblock layers, a plurality of activation layers, and one shuffle layer. Here, the at least two convolutional layers are respectively denoted as convolutional layer 0, convolutional layer 1, convolutional layer 2, convolutional layer 3, ……, convolutional layer N, and the activation layers include activation layer 0, activation layer 1, activation layer 2, ……. When the target encoding / decoding information includes the encoding / decoding information corresponding to the first related image block, the first related image block includes related image block 0, and the encoding / decoding information corresponding to the related image block 0 is encoding / decoding information 0. The computer device inputs the related image block 0 and the encoding / decoding information 0 into the dot product module, and the dot product module performs dot product operation processing on the related image block 0 and the encoding / decoding information 0 to obtain the fusion data corresponding to the related image block 0, inputs the fusion data corresponding to the related image block 0 into the convolutional layer 0, and the convolutional layer 0 performs convolutional operation processing on the fusion data corresponding to the related image block 0 to obtain the target fusion data corresponding to the first related image block.

[0076] Understandably, step S102 described above includes the following steps. The computer device obtains the image distance between the target image block and the first related image block from the target encoding / decoding information. If the image distance is greater than the distance threshold, it indicates that the similarity between the target image block and the first related image block is relatively low, that is, the significance of the target image block referring to the first related image block is low. In this case, based on the first related image block, filtering processing on the target image block is prohibited. If the image distance is less than or equal to the distance threshold, it indicates that the similarity between the target image block and the first related image block is relatively high, that is, the significance of the target image block referring to the first related image block is high. Therefore, the computer device can perform filtering processing on the target image block based on the first related image block to obtain a filtering image block corresponding to the target image block. By performing filtering processing on the target image block with an image distance less than or equal to the distance threshold, it is possible to narrow down the target and utilize the first related image block to improve the filtering performance of the image block.

[0077] Understandably, when determining the image distance between the first related image block and the target image block based on the image sequence count of the first related image block and the image sequence count of the target image block, the distance threshold is 2. If the absolute value corresponding to the difference between the image sequence count of the first related image block and the image sequence count of the target image block is less than or equal to 2, based on the first related image block and the target encoding / decoding information, filtering processing is performed on the target image block to obtain a filtering image block corresponding to the target image block. If the absolute value corresponding to the difference between the image sequence count of the first related image block and the image sequence count of the target image block is greater than 2, based on the first related image block, filtering processing on the target image block is prohibited to obtain a filtering image block corresponding to the target image block.

[0078] Understandably, when the image distance between the first related image block and the target image block is determined based on the time-domain layer of the target image block, the image distance between the first related image block and the target image block has an inverse correlation with the time-domain layer of the target image block. For example, when the distance threshold is 4 and the time-domain layer of the target image block is 4 or more, based on the first related image block and the target encoding / decoding information, filtering processing is performed on the target image block to obtain a filtering image block corresponding to the target image block. When the time-domain layer of the target image block is less than 4, it is prohibited to perform filtering processing on the target image block based on the first related image block to obtain a filtering image block corresponding to the target image block.

[0079] For example, as shown in FIG. 13e, an image filtering processor based on a neural network includes at least two convolutional layers, N Resblock layers, a plurality of activation layers, and one shuffle layer. Here, the at least two convolutional layers are respectively denoted as convolutional layer 0, convolutional layer 1, convolutional layer 2, convolutional layer 3, ……, convolutional layer N, and the activation layers include activation layer 0, activation layer 1, activation layer 2, ……. When the image distance between the first related image block and the target image block is greater than the distance threshold, the computer device can prohibit performing filtering processing on the target image block based on the first related image block. When the image distance between the first related image block and the target image block is less than or equal to the distance threshold, the computer device can input the target encoding / decoding information and the first related image block into the image filtering processor based on the neural network. Here, the target encoding / decoding information includes the encoding / decoding information corresponding to the target image block. Here, the first related image includes related image block 0 and related image block 1. That is, the convolutional layer 0 of the image filtering processor based on the neural network performs a convolutional operation on related image 0 to obtain the fusion data corresponding to related image block 0. The convolutional layer 1 of the image filtering processor based on the neural network performs a convolutional operation on related image 1 to obtain the fusion data corresponding to related image block 1, and determines the fusion data corresponding to related image block 0 and the fusion data corresponding to related image block 1 as the target fusion data corresponding to the related image block. In FIG. 13e, the encoding / decoding information corresponding to the target image block includes the sequence level quantization parameter, the slice level quantization parameter, the slice level encoding type, the block encoding type, and the predicted image block corresponding to the target image block.

[0080] In another example, as shown in FIG. 13f, an image filtering processor based on a neural network includes at least two convolutional layers, N Resblock layers, a plurality of activation layers, and one shuffle layer. Here, the at least two convolutional layers are respectively denoted as convolutional layer 0, convolutional layer 1, convolutional layer 2, convolutional layer 3, ……, convolutional layer N, and the activation layers include activation layer 0, activation layer 1, activation layer 2, ……. When the image distance between the first related image block and the target image block is greater than the distance threshold, the computer device can prohibit performing filtering processing on the target image block based on the first related image block. When the image distance between the first related image block and the target image block is less than or equal to the distance threshold, the computer device can input the target encoding / decoding information and the first related image block into the image filtering processor based on the neural network. Here, the target encoding / decoding information includes the encoding / decoding information corresponding to the target image block. Here, the first related images include related image block 0 and related image block 1. That is, the convolutional layer 0 of the image filtering processor based on the neural network performs convolutional operation processing on related image 0 and related image 1 to obtain target fusion data corresponding to the related image block. In FIG. 13f, the encoding / decoding information corresponding to the target image block includes a sequence level quantization parameter, a slice level quantization parameter, a slice level encoding type, a block encoding type, and a predicted image block corresponding to the target image block.

[0081] Understandably, the above step S102 includes the following steps. The computer device can obtain the image distance between the target image block and the first related image block from the target encoding / decoding information. When the image distance does not belong to the specified image distance, filtering processing on the target image block based on the first related image block is prohibited. When the image distance belongs to the specified image distance, filtering processing is performed on the target image block based on the first related image block and the target encoding / decoding information to obtain a filtering image block corresponding to the target image block. For example, the specified image distance may refer to that the time domain layer of the target image block is the specified time domain layer. For example, the specified time domain layer may be a layer that is a multiple of 2, Layer5, Layer4, and Layer5. Or the specified image distance may refer to that the difference between the image sequence count of the target image block and the image sequence count of the first related image block is the specified count difference. For example, the specified count difference may be 1, or 1 and 2, etc. By the image distance between the first related image block and the target image block, the target is narrowed down and the first related image block is used to improve the filtering performance of the image block.

[0082] Understandably, when the size of the target image block is smaller than the size of the first related image block, step S102 above includes the following steps. The computer device can divide the first related image block to obtain N (N is a positive integer greater than 1) image sub-blocks, and the target encoding / decoding information includes encoding / decoding information for representing the influence degree of each of the N image sub-blocks on the target image block. The sizes of the N image sub-blocks may be the same or different. For example, as shown in FIG. 8b, the computer device can divide the first related image block into 5 image sub-blocks. The 5 image sub-blocks include the gray area in the reference image and 4 hatched areas, and the sizes of the 5 image sub-blocks are the same. Further, the N image sub-blocks and the encoding / decoding information corresponding to each of the N image sub-blocks are input into an image filtering processor based on a neural network. The computer device can perform filtering processing on the target image block based on the N image sub-blocks and the encoding / decoding information corresponding to each of the N image sub-blocks by the image filtering processor based on the neural network, and obtain a filtering image block corresponding to the target image block. By introducing related image blocks in a wider area, more information is provided in the encoding process of the target image block, and the accuracy of the filtering process for the target image is improved.

[0083] Understandably, after obtaining the filtering image block corresponding to the target image block, the computer device can generate a second related image block for decoding the predicted image block in the multimedia data based on the filtering image block corresponding to the target image block, and store the second related image block in the decoding cache. Alternatively, the computer device can generate a display image block based on the filtering image block corresponding to the target image block, skip the step of storing the filtering image block corresponding to the target image block in the decoding cache, and send the display image block to the display device, and the display device is configured to display the display image block.

[0084] In a specific implementation, as shown in FIG. 13g, an image filtering processor based on a neural network includes at least two convolutional layers, N Resblock layers, a plurality of activation layers, and one shuffle layer. Here, the at least two convolutional layers are respectively denoted as convolutional layer 0, convolutional layer 1, convolutional layer 2, convolutional layer 3, ……, convolutional layer N, and the activation layers include activation layer 0, activation layer 1, activation layer 2, ……. When the target encoding / decoding information includes the encoding / decoding information corresponding to the first related image block and the encoding / decoding information corresponding to the target image block, the first related image block includes related image block 0 and related image block 1. The encoding / decoding information corresponding to the related image block 0 is encoding / decoding information 0, and the encoding / decoding information corresponding to the related image block 1 is encoding / decoding information 1. The encoding / decoding information 1 includes the image distance 1 between the related image 1 and the target image block and the reference direction 1 of the related image 1 with respect to the target image block. The encoding / decoding information 0 includes the image distance 0 between the related image 0 and the target image block and the reference direction 0 of the related image 0 with respect to the target image block. The convolutional layer 0 corresponds to the related image block 0, and the convolutional layer 1 corresponds to the related image block 1. The encoding / decoding information corresponding to the target image block includes a sequence level quantization parameter, a slice level quantization parameter, a slice level encoding type, a block encoding type, and a predicted image block corresponding to the target image block. The computer device can input the encoding / decoding information corresponding to the target image block as auxiliary information into the image filtering processor based on the neural network together with the encoding / decoding information of the related image block. FIG. 13g shows an example, adopting a matrix to represent each of the related image block 0, the reference direction 0, and the image distance 0. The convolutional layer 0 performs a convolution operation process on the matrix corresponding to the related image block 0, the matrix corresponding to the reference direction 0, and the matrix corresponding to the image distance 0 to obtain the fusion data corresponding to the related image block 0.Adopt a matrix to represent each of the associated image block 1, reference direction 1, and image distance 1, and perform convolution operation processing on the matrix corresponding to the associated image block 1, the matrix corresponding to the reference direction 1, and the matrix corresponding to the image distance 1 by the convolution layer 1 to obtain the fusion data corresponding to the associated image block 1. Determine the fusion data corresponding to the associated image block 1 and the fusion data corresponding to the associated image block 0 as the target fusion data corresponding to the associated image block. At the same time, adopt a matrix to represent each of the sequence-level quantization parameter, slice-level quantization parameter, slice-level coding type, block coding type, and the predicted image block corresponding to the target image block, input it into the image filtering processor based on a neural network, and perform filtering processing on the target image block by the image filtering processor based on the above target fusion data and target encoding / decoding information to obtain the filtered image block corresponding to the target image block.

[0085] Understandably, when the image to which the target image block belongs is an I-frame, a matrix with all elements being 0 is determined as the matrix corresponding to the slice encoding type corresponding to the target image block. When the image to which the target image block belongs is a B-frame or a P-frame, a matrix with all elements being 1 is determined as the matrix corresponding to the slice encoding type corresponding to the target image block. The matrix corresponding to the block encoding type of the target image block can be filled element by element according to the encoding method of the encoding unit. For example, when the encoding unit belongs to the intra encoding method, the value of the element corresponding to the encoding unit is set to 0, and otherwise, the value is set to 1. When the image to which the target image block belongs is an I-frame, the related image block 0 and the related image 1 can be filled using the target image block. The image distance 1 and the image distance 0 are set to the minimum weight 0, and the matrix corresponding to the image distance 1 and the matrix corresponding to the image distance 0 can be obtained. When the image to which the target image block belongs is a B-frame or a P-frame, the image distance 1 is set to (1025 - the absolute value of the POC difference between the related image block 1 and the target image block) / 1024 to obtain the matrix corresponding to the image distance 1, and the image distance 0 is set to (1025 - the absolute value of the POC difference between the related image block 0 and the target image block) / 1024 to obtain the matrix corresponding to the image distance 0. When the image to which the target image block belongs is an I-frame, the direction values (i.e., the elements of the matrix) corresponding to the reference direction 0 and the reference direction 1 are both set to 0. When the image to which the target image block belongs is a B-frame or a P-frame and the POC of the image to which the related image block 0 belongs is smaller than the POC of the image to which the target image block belongs, the direction value of the reference direction 0 is set to 0, and otherwise, the direction value of the reference direction 0 is set to 1. When the POC of the image to which the related image block 1 belongs is smaller than the POC of the image to which the target image block belongs, the direction value of the reference direction 1 is set to 0, and otherwise, the direction value of the reference direction 1 is set to 1.When the image to which the target image block belongs is a B-frame or a P-frame, the related image block 0 may be obtained from the first reference frame in the reference frame list in the L0 direction, and the related image block 1 may be obtained from the first reference frame in the reference frame list in the L1 direction. The related image block 0 and the related image block 1 both have the same spatial position relationship as the target image block, and the size of the related image block 0 and the size of the related image block 1 are both the same as the size of the target image block. By the above processing method, it is possible to enable the I-frame, B-frame, and P-frame to share the above neural network-based image filtering processor for filtering processing, without the need to train a neural network-based image filtering processor for each of the I-frame, B-frame, and P-frame, saving resources and improving the versatility of the neural network-based image filtering processor.

[0086] In the embodiments of the present application, a first related image block related to a target image block to be filtered in multimedia data is determined, the first related image block and target encoding / decoding information related to the target image block are obtained, and based on the target encoding / decoding information and the first related image block, filtering processing is performed on the target image block to obtain a filtering image block corresponding to the target image block. By introducing multi-dimensional information such as related image blocks and target encoding / decoding information, a rich amount of information is provided in the filtering process of the target image block. At the same time, since the target encoding / decoding information is used to represent the degree of influence of the first related image block on the target image block, based on the target encoding / decoding information and an image filtering processor based on a neural network, the high-quality advantages of the related image blocks are fully utilized to perform filtering processing on the reconstructed image block. In the filtering processing process for the target image block, the related image blocks are differentially used to improve the accuracy of the filtering processing of the target image block and help improve the quality of the filtering processing of the target image block, that is, the accuracy of the image filtering processing can be improved and the filtering effect of the image can be improved. Furthermore, the encoding quality of the multimedia data can be improved.

[0087] Referring to FIG. 14, it is an exemplary structural diagram of a multimedia data processing apparatus according to an embodiment of the present application. The above multimedia data processing apparatus may be a computer program (including program code) executed by computer equipment. For example, the multimedia data processing apparatus may be an application software, and the apparatus may be configured to execute corresponding steps in the method provided by the embodiments of the present application. As shown in FIG. 14, the multimedia data processing apparatus may include an acquisition module 141, a filtering module 142, a determination module 143, and a generation module 144.

[0088] The decision module is configured to determine a first related image block related to a target image block to be filtered in multimedia data, the acquisition module is configured to acquire target encoding / decoding information related to the target image block, and the filtering module is configured to perform a filtering process on the target image block by an image filtering processor based on a neural network according to the first related image block and the target encoding / decoding information, so as to obtain a filtering image block corresponding to the target image block.

[0089] Understandably, the filtering module includes a fusion unit 14a and a filtering unit 15a. The fusion unit is configured to fuse the first related image block with encoding / decoding information corresponding to the first related image block by an information fusion layer of an image filtering processor based on a neural network, so as to obtain target fusion data corresponding to the first related image block. The filtering unit is configured to perform a filtering process on the target image block according to the target fusion data corresponding to the first related image block by a filtering process layer of the image filtering processor based on the neural network, so as to obtain a filtering image block corresponding to the target image block.

[0090] Understandably, the fusion unit is configured to fuse the first related image block with encoding / decoding information corresponding to the first related image block by an information fusion layer of an image filtering processor based on a neural network, so as to obtain target fusion data corresponding to the first related image block. The filtering unit is configured to perform a filtering process on the target image block according to the target fusion data corresponding to the first related image block and the encoding / decoding information corresponding to the target image block by a filtering process layer of the image filtering processor based on the neural network, so as to obtain a filtering image block corresponding to the target image block.

[0091] It is understandable that the number of the first related image blocks is M (M is an integer greater than or equal to 1), the image filtering processor based on the neural network includes M information fusion layers, one information fusion layer corresponds to one related image block, and the fusion unit fuses the first related image block with the encoding / decoding information corresponding to the first related image block by the information fusion layer of the image filtering processor based on the neural network to obtain the target fusion data corresponding to the first related image block. This includes performing a convolution operation process on the first related image block i (i is a positive integer less than or equal to M) and the encoding / decoding information corresponding to the first related image block i by the information fusion layer of the image filtering processor based on the neural network to obtain the fusion data corresponding to the first related image block i. The first related image block i belongs to the M first related image blocks, and when the fusion data corresponding to each of the M first related image blocks is obtained, determining the fusion data corresponding to each of the M first related image blocks as the target fusion data corresponding to the first related image block.

[0092] It is understandable that here, the number of the first related image blocks is M (M is an integer greater than or equal to 1), and the fusion unit fuses the first related image block with the encoding / decoding information corresponding to the first related image block by the information fusion layer of the image filtering processor based on the neural network to obtain the target fusion data corresponding to the first related image block. This includes performing a convolution operation process on the M first related image blocks and the encoding / decoding information by the information fusion layer of the image filtering processor based on the neural network to obtain the target fusion data corresponding to the first related image block.

[0093] It is understandable that the number of the first related image blocks is M (M is an integer greater than or equal to 1), and the fusion unit fuses the first related image blocks with the encoding / decoding information corresponding to the first related image blocks through the information fusion layer of the image filtering processor based on a neural network to obtain the target fusion data corresponding to the first related image blocks. This is to perform a dot product operation on the first related image block i (i is a positive integer less than or equal to M) and the encoding / decoding information corresponding to the first related image block i through the information fusion layer of the image filtering processor based on a neural network to obtain the fusion data corresponding to the first related image block i. The first related image block i belongs to the M first related image blocks. When the fusion data corresponding to each of the M first related image blocks is obtained, a convolution operation is performed on the fusion data corresponding to each of the M first related image blocks to obtain the target fusion data corresponding to the first related image block, and it includes the above.

[0094] It is understandable that the encoding / decoding information corresponding to the target image block includes at least one of a sequence level quantization parameter, a slice level quantization parameter, a slice level encoding type, a block encoding type, a filtering enhanced image corresponding to the target image block, prediction image block information corresponding to the target image block, and block division image block information of the target image block. It is understandable that the image block information corresponding to the target image block includes at least one of first color component information, second color component information, and third color component information. Here, the first color component information is used to represent the luminance of the image block, and the second color component information and the third color component information are used to represent the chroma of the image block. The image block information corresponding to the target image block includes any one of filtering intensity image block information, block division image block information, and prediction image block information corresponding to the target image block.

[0095] Understandably, the target encoding / decoding information is used to represent the degree of influence of the first related image block on the target image block, and the target encoding / decoding information includes at least one of the influence factor of the first related image block on the target image block, the reference direction of the first related image block with respect to the target image block, the block split image block information of the first related image block, the quantization parameter corresponding to the first related image block, the reconstructed image block information corresponding to the first related image block, the filtering intensity image block information corresponding to the first related image block, and the predicted image block information corresponding to the first related image block. Here, the influence factor is determined based on the image distance between the first related image block and the target image.

[0096] When the slice where the target image block is located is an incomplete intra-coded slice, the determination module determines the image distance between the first related image block and the target image block based on the image order count corresponding to the first related image block and the image order count corresponding to the target image block. When the slice where the target image block is located is a complete intra-coded slice, the determination module is configured to determine a predetermined value as the image distance between the first related image block and the target image block.

[0097] Understandably, the determination module determines the image distance between the first related image block and the target image block based on the image order count corresponding to the first related image block and the image order count corresponding to the target image block, which includes determining the count difference between the image order count corresponding to the first related image block and the image order count corresponding to the target image block, and determining the absolute value of the count difference as the image distance between the first related image block and the target image block.

[0098] Understandably, the determining module determines the image distance between the first related image block and the target image block based on the image sequence count corresponding to the first related image block and the image sequence count corresponding to the target image block, which includes determining the count difference between the image sequence count corresponding to the first related image block and the image sequence count corresponding to the target image block, calculating a count mapping value corresponding to the absolute value of the count difference, and determining the image distance between the first related image block and the target image based on the count mapping value.

[0099] Understandably, the image distance between the first related image block and the target image block is determined based on the time domain layer of the target image block, or the image distance between the first related image block and the target image block is determined based on a level mapping value corresponding to the time domain layer of the target image block.

[0100] Understandably, the image block information corresponding to the first related image block includes at least one of first color component information, second color component information, and third color component information. Here, the first color component information is used to represent the luminance of the image block, the second color component information and the third color component information are used to represent the chroma of the image block, and the image block information corresponding to the first related image block includes any one of the reconstructed image block information, block split image block information, filtering intensity image block information, and predicted image block information corresponding to the first related image block.

[0101] It is understandable that when the slice where the target image block is located is an incomplete intra-coded slice, the determination module further determines the reference direction of the first related image block with respect to the target image block based on the magnitude relationship between the picture order count corresponding to the first related image block and the picture order count corresponding to the target image block. When the slice where the target image block is located is a complete intra-coded slice, it is configured to adopt a predetermined direction value to indicate the reference direction of the first related image block with respect to the target image block.

[0102] It is understandable that the filtering module performing filtering processing on the target image block by an image filtering processor based on a neural network based on the first related image block and the target coding / decoding information to obtain a filtered image block corresponding to the target image block includes: obtaining an image distance between the target image block and the first related image block from the target coding / decoding information; prohibiting performing filtering processing on the target image block based on the first related image block when the image distance is greater than a distance threshold; and when the image distance is less than or equal to the distance threshold, performing filtering processing on the target image block by an image filtering processor based on a neural network based on the first related image block and the target coding / decoding information to obtain a filtered image block corresponding to the target image block.

[0103] It is understandable that based on the filtering module, the first related image block, and the target encoding / decoding information, an image filtering processor based on a neural network performs a filtering process on the target image block to obtain a filtered image block corresponding to the target image block, which includes obtaining an image distance between the target image block and the first related image block from the target encoding / decoding information, and when the image distance does not belong to a specified image distance, prohibiting performing a filtering process on the target image block based on the first related image block, and when the image distance belongs to the specified image distance, performing a filtering process on the target image block by an image filtering processor based on a neural network based on the first related image block and the target encoding / decoding information to obtain a filtered image block corresponding to the target image block.

[0104] It is understandable that the size of the target image block is the same as the size of the first related image block, or the size of the target image block is different from the size of the first related image block. There is the same spatial position relationship between the target image block and the first related image block.

[0105] It is understandable that the size of the target image block is the same as the size of the first related image block, or the size of the target image block is different from the size of the first related image block. Here, there is a different spatial position relationship between the target image block and the first related image block, and the first related image block is determined based on the motion vector of the target image block.

[0106] Understandably, when the size of the target image block is smaller than the size of the first related image block by the filtering module, based on the first related image block and the target encoding / decoding information, performing filtering processing on the target image block to obtain a filtering image block corresponding to the target image block means dividing the first related image block to obtain N (N is a positive integer greater than 1) image sub-blocks, where the target encoding / decoding information includes encoding / decoding information for representing the influence degree of each of the N image sub-blocks on the target image block, inputting the N image sub-blocks and the encoding / decoding information corresponding to each of the N image sub-blocks into an image filtering processor based on a neural network, and performing filtering processing on the target image block based on the N image sub-blocks and the encoding / decoding information corresponding to each of the N image sub-blocks by the image filtering processor based on the neural network to obtain a filtering image block corresponding to the target image block.

[0107] Understandably, the generation module is configured to generate a second related image block for decoding a predicted image block in the multimedia data based on the filtering image block corresponding to the target image block and store the second related image block in a decoding cache, or generate a display image block based on the filtering image block corresponding to the target image block, skip the step of storing the filtering image block corresponding to the target image block in the decoding cache, and be configured to transmit the display image block to a display device, and the display device is configured to display the display image block.

[0108] In an embodiment of the present application, a first related image block related to a target image block to be filtered in multimedia data is determined, target encoding / decoding information of the first related image block and the target image block is obtained, and based on the target encoding / decoding information and the first related image block, filtering processing is performed on the target image block to obtain a filtering image block corresponding to the target image block. By introducing multi-dimensional information such as related image blocks and target encoding / decoding information, while providing a rich amount of information through the filtering process of the target image block, since the target encoding / decoding information is used to represent the degree of influence of the first related image block on the target image block, the high-quality advantages of the related image blocks are fully utilized by the target encoding / decoding information to perform filtering processing on the reconstructed image block, and in the filtering processing process for the target image block, the related image blocks are differentially used to improve the accuracy of the filtering processing of the target image block, which helps to improve the quality of the filtering processing of the target image block, that is, it can improve the accuracy of the filtering processing of the image and the filtering effect of the image, and further, the encoding quality of the multimedia data can be improved.

[0109] Referring to FIG. 15, it is an exemplary structural diagram of a computer device according to an embodiment of the present application. As shown in FIG. 15, the computer device 1000 described above can include a processor 1001, a network interface 1004, and a memory 1005. In addition, the computer device 1000 described above can further include a user interface 1003 and at least one communication bus 1002. Here, the communication bus 1002 is configured to realize connection communication between these components. Here, the user interface 1003 can include a display and a keyboard. Optionally, the user interface 1003 can further include a standard wired interface and a wireless interface. The network interface 1004 can optionally include a standard wired interface and a wireless interface (e.g., a WI-FI interface). The memory 1005 can be a high-speed RAM memory, or a non-volatile memory, such as at least one magnetic disk memory, for example. Optionally, the memory 1005 can further be at least one storage device separated from the aforementioned processor 1001. As shown in FIG. 15, the memory 1005 as a computer-readable storage medium can include an operating system, a network communication module, a network interface module, and a device control application program.

[0110] In the computer device 1000 shown in FIG. 15, the network interface 1004 can provide a network communication function, and the user interface 1003 is mainly configured to provide an interface for input to media content.

[0111] Understandably, the processor 1001 calls the device-limited application program stored in the memory 1005, Determining a first related image block related to a target image block to be filtered within multimedia data; Obtaining target encoding / decoding information related to the first related image block and the target image block; Based on the first related image block and the target encoding / decoding information, performing a filtering process on the target image block by an image filtering processor based on a neural network to obtain a filtering image block corresponding to the target image block, and can be configured to implement.

[0112] Understandably, the processor 1001 can further call a device-limited application program stored in the memory 1005, and based on the first related image block and the target encoding / decoding information, perform a filtering process on the target image block by an image filtering processor based on a neural network to obtain a filtering image block corresponding to the target image block, and the step is Fusing, by an information fusion layer of an image filtering processor based on a neural network, the first related image block with encoding / decoding information corresponding to the first related image block to obtain target fusion data corresponding to the first related image block; Based on the target fusion data corresponding to the first related image block, performing a filtering process on the target image block by a filtering process layer of the image filtering processor based on a neural network to obtain a filtering image block corresponding to the target image block, and includes.

[0113] Understandably, the processor 1001 can be configured to call the device-limited application program stored in the memory 1005, and based on the first related image block and the target encoding / decoding information, perform filtering processing on the target image block by an image filtering processor based on a neural network to obtain a filtering image block corresponding to the target image block. The steps are as follows: A step of fusing the first related image block with the encoding / decoding information corresponding to the first related image block by an information fusion layer of an image filtering processor based on a neural network to obtain target fusion data corresponding to the first related image block; A step of performing filtering processing on the target image block based on the target fusion data corresponding to the first related image block and the encoding / decoding information corresponding to the target image block by a filtering processing layer of an image filtering processor based on a neural network to obtain a filtering image block corresponding to the target image block. The steps include:

[0114] Understandably, the number of the first related image blocks is M (M is an integer greater than or equal to 1), and the image filtering processor based on a neural network includes M information fusion layers. One information fusion layer corresponds to one related image block. The processor 1001 can be configured to call the device-limited application program stored in the memory 1005, and a step of fusing the first related image block with the encoding / decoding information corresponding to the first related image block by an information fusion layer of an image filtering processor based on a neural network to obtain target fusion data corresponding to the first related image block. The steps are as follows: The step of performing a convolution operation on the first related image block i (where i is a positive integer not exceeding M) and the encoding / decoding information corresponding to the first related image block i by the information fusion layer i of the image filtering processor based on the neural network to obtain the fusion data corresponding to the first related image block i, where the first related image block i belongs to M first related image blocks, and the information fusion layer i is the information fusion layer corresponding to the first related image block i among the M information fusion layers, When the fusion data corresponding to each of the M first related image blocks is obtained, determining the fusion data corresponding to each of the M first related image blocks as the target fusion data corresponding to the first related image block,

[0115] It can be understood that the number of the first related image blocks is M (M is an integer greater than or equal to 1), The processor 1001 can be configured to realize the step of calling the device-limited application program stored in the memory 1005 and fusing the first related image block with the encoding / decoding information corresponding to the first related image block by the information fusion layer of the image filtering processor based on the neural network to obtain the target fusion data corresponding to the first related image block. The step includes The step of performing a convolution operation on M first related image blocks and encoding / decoding information by the information fusion layer of the image filtering processor based on the neural network to obtain the target fusion data corresponding to the first related image block.

[0116] It can be understood that the number of the first related image blocks is M (M is an integer greater than or equal to 1), and the image filtering processor based on the neural network includes M information fusion layers, and one information fusion layer corresponds to one related image block. The processor 1001 is configured to call the device-limited application program stored in the memory 1005, and fuse the first related image block with the encoding / decoding information corresponding to the first related image block by the information fusion layer of the neural network-based image filtering processor, so as to obtain the target fusion data corresponding to the first related image block. The steps are as follows: The information fusion layer i of the neural network-based image filtering processor performs a dot product operation on the first related image block i (where i is a positive integer not exceeding M) and the encoding / decoding information corresponding to the first related image block i, to obtain the fusion data corresponding to the first related image block i. The first related image block i belongs to M first related image blocks, and the information fusion layer i is the information fusion layer corresponding to the first related image block i among the M information fusion layers. When the fusion data corresponding to each of the M first related image blocks is obtained, a convolution operation is performed on the fusion data corresponding to each of the M first related image blocks, to obtain the target fusion data corresponding to the first related image block.

[0117] It can be understood that the encoding / decoding information corresponding to the target image block includes at least one of the sequence level quantization parameter, the slice level quantization parameter, the slice level encoding type, the block encoding type, the filtering intensity image block information corresponding to the target image block, the predicted image block information corresponding to the target image block, and the block segmentation image block information of the target image block.

[0118] Understandably, the image block information corresponding to the target image block includes at least one of first color component information, second color component information, and third color component information. Here, the first color component information is used to represent the luminance of the image block, the second color component information and the third color component information are used to represent the chroma of the image block, and the image block information corresponding to the target image block includes any one of the filtering intensity image block information, block division image block information, and prediction image block information corresponding to the target image block.

[0119] Understandably, the target encoding / decoding information is used to represent the influence degree of the first related image block on the target image block. The target encoding / decoding information includes at least one of the influence factor of the first related image block on the target image block, the reference direction of the first related image block with respect to the target image block, the block division image block information of the first related image block, the quantization parameter corresponding to the first related image block, the reconstructed image block information corresponding to the first related image block, the filtering intensity image block information corresponding to the first related image block, and the prediction image block information corresponding to the first related image block. Here, the influence factor is determined based on the image distance between the first related image block and the target image block.

[0120] Understandably, when the slice where the target image block is located is an incomplete intra-coded slice, the processor 1001 further calls the device-limited application program stored in the memory 1005, and based on the image order count corresponding to the first related image block and the image order count corresponding to the target image block, determines the image distance between the first related image block and the target image block. When the slice in which the target image block is located is a complete intra-coded slice, a step of determining a predetermined distance value as the image distance between the first related image block and the target image block can be configured to be realized.

[0121] Understandably, the processor 1001 can be configured to call a device-limited application program stored in the memory 1005 and determine the image distance between the first related image block and the target image block based on the image sequence count corresponding to the first related image block and the image sequence count corresponding to the target image block. The step includes: a step of determining a count difference between the image sequence count corresponding to the first related image block and the image sequence count corresponding to the target image block; a step of determining the absolute value of the count difference as the image distance between the first related image block and the target image block.

[0122] Understandably, the processor 1001 can be configured to call a device-limited application program stored in the memory 1005 and determine the image distance between the first related image block and the target image block based on the image sequence count corresponding to the first related image block and the image sequence count corresponding to the target image block. The step includes: a step of determining a count difference between the image sequence count corresponding to the first related image block and the image sequence count corresponding to the target image block; a step of calculating a count mapping value corresponding to the absolute value of the count difference and determining the image distance between the first related image block and the target image block based on the count mapping value.

[0123] It is understandable that the image distance between the first related image block and the target image block is determined based on the temporal region layer of the target image block, or the image distance between the first related image block and the target image block is determined based on a level mapping value corresponding to the temporal region layer of the target image block.

[0124] It is understandable that the image block information corresponding to the first related image block includes at least one of first color component information, second color component information, and third color component information. Here, the first color component information is used to represent the luminance of the image block, the second color component information and the third color component information are used to represent the chroma of the image block, and the image block information corresponding to the first related image block includes any one of the reconstructed image block information, block split image block information, filtering intensity image block information, and predicted image block information corresponding to the first related image block.

[0125] It is understandable that when the processor 1001 calls the device-limited application program stored in the memory 1005 and the slice where the target image block is located is an incomplete intra-coded slice, based on the relationship between the image order count corresponding to the first related image block and the image order count corresponding to the target image block, determining the reference direction of the first related image block with respect to the target image block; When the slice where the target image block is located is a complete intra-coded slice, adopting a predetermined direction value to indicate the reference direction of the first related image block with respect to the target image block. It can be configured to implement the above steps.

[0126] Understandably, the processor 1001 can be configured to call the device-limited application program stored in the memory 1005, and based on the first related image block and the target encoding / decoding information, perform filtering processing on the target image block by an image filtering processor based on a neural network to obtain a filtering image block corresponding to the target image block. The steps are as follows: Obtaining an image distance between the target image block and the first related image block from the target encoding / decoding information; When the image distance is greater than a distance threshold, prohibiting performing filtering processing on the target image block based on the first related image block; When the image distance is less than or equal to the distance threshold, performing filtering processing on the target image block by an image filtering processor based on a neural network based on the first related image block and the target encoding / decoding information to obtain a filtering image block corresponding to the target image block.

[0127] Understandably, the processor 1001 can be configured to call the device-limited application program stored in the memory 1005, and based on the first related image block and the target encoding / decoding information, perform filtering processing on the target image block by an image filtering processor based on a neural network to obtain a filtering image block corresponding to the target image block. The steps are as follows: Obtaining an image distance between the target image block and the first related image block from the target encoding / decoding information; When the image distance does not belong to a specified image distance, prohibiting performing filtering processing on the target image block based on the first related image block; When the image distance belongs to the specified image distance, based on the first related image block and the target encoding / decoding information, an image filtering processor based on a neural network performs filtering processing on the target image block to obtain a filtered image block corresponding to the target image block, including the step of:

[0128] It is understandable that the size of the target image block is the same as the size of the first related image block, or the size of the target image block is different from the size of the first related image block. There is the same spatial position relationship between the target image block and the first related image block.

[0129] It is understandable that the size of the target image block is the same as the size of the first related image block, or the size of the target image block is different from the size of the first related image block. There is a different spatial position relationship between the target image block and the first related image block, and the first related image block is determined based on the motion vector of the target image block.

[0130] It is understandable that the processor 1001 calls a device-limited application program stored in the memory 1005. When the size of the target image block is smaller than the size of the first related image block, based on the first related image block and the target encoding / decoding information, an image filtering processor based on a neural network performs filtering processing on the target image block to obtain a filtered image block corresponding to the target image block. The step can be configured to be realized, and the step includes: The step of dividing the first related image block to obtain N (N is a positive integer greater than 1) image sub-blocks, where the target encoding / decoding information includes encoding / decoding information for representing the influence degree of each of the N image sub-blocks on the target image block. Inputting the N image sub-blocks and the encoding / decoding information corresponding to each of the N image sub-blocks into an image filtering processor based on a neural network; Performing a filtering process on the target image block by the image filtering processor based on the neural network according to the N image sub-blocks and the encoding / decoding information corresponding to each of the N image sub-blocks, to obtain a filtered image block corresponding to the target image block. The method includes the steps above.

[0131] Understandably, the processor 1001 calls a device-limited application program stored in the memory 1005 to Generate a second related image block for decoding a predicted image block in the multimedia data based on the filtered image block corresponding to the target image block, and store the second related image block in a decoding cache, or Skip the step of generating a display image block based on the filtered image block corresponding to the target image block and storing the filtered image block corresponding to the target image block in a decoding cache, and be configured to implement the step of transmitting the display image block to a display device, where the display device is configured to display the display image block.

[0132] In an embodiment of the present application, a first related image block related to a target image block to be filtered in multimedia data is determined, target encoding / decoding information of the influence degree of the first related image block and the target image block is obtained, and based on the target encoding / decoding information and the first related image block, filtering processing is performed on the target image block to obtain a filtering image block corresponding to the target image block. By introducing multi-dimensional information such as related image blocks and target encoding / decoding information, while providing a rich amount of information in the filtering process of the target image block, since the target encoding / decoding information is used to represent the influence degree of the first related image block on the target image block, based on the target encoding / decoding information and an image filtering processor based on a neural network, the high-quality advantages of the related image blocks are fully utilized to perform filtering processing on the reconstructed image blocks, and in the filtering processing process for the target image block, the related image blocks are differentially used to improve the accuracy of the filtering processing of the target image block, which helps to improve the quality of the filtering processing of the target image block, that is, the accuracy of the image filtering processing can be improved, the filtering effect of the image can be improved, and furthermore, the encoding quality of the multimedia data can be improved.

[0133] It should be understood that the computer device 1000 described in the embodiment of the present application can execute the above description of the multimedia data processing method in the embodiment corresponding to FIG. 4 described above, and can also execute the above description of the multimedia data processing apparatus in the embodiment corresponding to FIG. 14 described above, which will not be repeated here. Furthermore, the beneficial effects of adopting the same method will not be repeated either.

[0134] Furthermore, it should be noted that the embodiments of the present application further provide a computer-readable storage medium, and the computer-readable storage medium stores a computer program executed by the multimedia data processing device mentioned in the above specification. The computer program includes program instructions. When the above processor executes the program instructions, it can execute the description of the above multimedia data processing method corresponding to FIG. 4 described above, which will not be repeated here. Furthermore, the beneficial effects of adopting the same method will not be repeated either. For the technical details not disclosed in the embodiments of the computer-readable storage medium according to the present application, please refer to the description of the method embodiments of the present application.

[0135] As an example, the above program instructions may be deployed to be executed on one computer device, or may be deployed to be executed on at least two computer devices in the same location, or may be deployed to be executed on at least two computer devices distributed in at least two locations and interconnected by a communication network. At least two computer devices distributed at least two locations and connected to each other via a communication network can constitute a blockchain network.

[0136] The above computer-readable storage medium may be an internal storage unit of the above computer device such as the hardware or memory of the data processing device or computer equipment according to any of the foregoing embodiments. The computer-readable storage medium may also be an external storage device of the computer device, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. implemented in the computer device. Further, the computer-readable storage medium may further include both the internal storage unit and the external storage device of the computer device. The computer-readable storage medium is configured to store the computer program and other programs and data required by the computer device. The computer-readable storage medium can further be configured to temporarily store the output or to-be-output data.

[0137] Embodiments of the present application further provide a computer program product including a computer program / instructions. When the computer program / instructions are executed by a processor, the processor is caused to implement the description of the above multimedia data processing method of the embodiment corresponding to FIG. 4 described above, and thus, it will not be repeatedly described herein. Further, the beneficial effects of adopting the same method will not be repeatedly described either. For technical details not disclosed in the embodiments of the computer program product according to the present application, reference may be made to the description of the method embodiments of the present application.

[0138] In the description, claims, and attached drawings of the embodiments of this application, terms such as "first", "second", and "third" do not limit a specific order, but are used to distinguish different media contents. Also, "including" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, apparatus, product, or device including a series of steps or units is not limited to the listed steps or modules, but further includes, by way of illustration, steps or modules not listed, or other step units specific to these processes, methods, apparatus, products, or devices also by way of illustration.

[0139] It will be apparent to those skilled in the art that the units and algorithm steps of each embodiment described with reference to the embodiments disclosed in this specification can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, in the above description, the configurations and steps of each example are described in general terms according to functions. Whether these functions are executed in the form of hardware or software depends on the specific application of the technical solution and the design constraints. Those skilled in the art can use different methods according to each specific application to realize the described functions, but such realizations should not be considered as exceeding the protection scope of this application.

[0140] The method and related apparatus provided by the embodiments of this application are described with reference to the flowchart of the method and / or exemplary structural diagrams according to the embodiments of this application. Specifically, by computer program instructions, each process and / or block of the flowchart of the method and / or exemplary structural diagrams, and combinations of processes and / or blocks within the flowchart and / or block diagrams may be realized. To generate a machine, by providing these computer program instructions to the processor of a general-purpose computer, a dedicated computer, an embedded processor, or other programmable data processing device, instructions executed by the processor of the computer or other programmable data processing device are caused to generate an apparatus for performing the functions specified in one or more processes of the flowchart and / or one or more blocks of the exemplary structural diagrams. These computer program instructions may also be stored in a computer-readable memory that causes the computer or other programmable data processing device to operate in a specific manner, and a product including instruction means for embodying the functions specified in one or more processes of the flowchart and / or one or more blocks of the exemplary structural diagrams is generated by the instructions stored in the computer-readable memory. These computer program instructions may be loaded into a computer or other programmable data processing device to cause the computer or other programmable device to execute a series of operation steps to generate a process implemented by the computer, whereby the instructions executed by the computer or other programmable device provide steps for embodying the functions specified in one or more processes of the flowchart and / or one or more blocks of the exemplary structural diagrams.

[0141] What is disclosed above is only the preferred embodiments of this application and cannot limit the claims of this application thereby. Therefore, the same changes made in accordance with the claims of this application are still included in the scope of this application.

Claims

1. A multimedia data processing method executed by a computer device, comprising: determining a first related image block related to a target image block to be filtered in multimedia data; obtaining target encoding / decoding information related to the target image block; performing a filtering process on the target image block by an image filtering processor based on a neural network according to the first related image block and the target encoding / decoding information, to obtain a filtered image block corresponding to the target image block. A multimedia data processing method.

2. The step of performing a filtering process on the target image block by an image filtering processor based on a neural network according to the first related image block and the target encoding / decoding information, to obtain a filtered image block corresponding to the target image block, comprises: obtaining, by an information fusion layer of the image filtering processor based on the neural network, target fusion data corresponding to the first related image block by fusing the first related image block with encoding / decoding information corresponding to the first related image block; performing a filtering process on the target image block according to the target fusion data corresponding to the first related image block by a filtering process layer of the image filtering processor based on the neural network, to obtain a filtered image block corresponding to the target image block. The multimedia data processing method according to Claim 1.

3. The step of performing a filtering process on the target image block by an image filtering processor based on a neural network according to the first related image block and the target encoding / decoding information, to obtain a filtered image block corresponding to the target image block, comprises: obtaining, by an information fusion layer of the image filtering processor based on the neural network, target fusion data corresponding to the first related image block by fusing the first related image block with encoding / decoding information corresponding to the first related image block; The filtering processing layer of the image filtering processor based on the neural network performs filtering processing on the target image block based on the target fusion data corresponding to the first related image block and the encoding / decoding information corresponding to the target image block, and obtains a filtered image block corresponding to the target image block. The multimedia data processing method according to claim 1.

4. The number of the first related image blocks is M, where M is an integer greater than or equal to 1. The image filtering processor based on the neural network includes M information fusion layers, and one information fusion layer corresponds to one related image block. The step of obtaining the target fusion data corresponding to the first related image block by fusing the first related image block with the encoding / decoding information corresponding to the first related image block by the information fusion layer of the image filtering processor based on the neural network is as follows: The step of performing a convolution operation on the first related image block i and the encoding / decoding information corresponding to the first related image block i by the information fusion layer i of the image filtering processor based on the neural network to obtain the fusion data corresponding to the first related image block i, where the first related image block i belongs to M first related image blocks, i is a positive integer less than or equal to M, and the information fusion layer i is the information fusion layer corresponding to the first related image block i among the M information fusion layers. When the fusion data corresponding to each of the M first related image blocks is obtained, determining the fusion data corresponding to each of the M first related image blocks as the target fusion data corresponding to the first related image block. The multimedia data processing method according to claim 2.

5. The number of the first related image blocks is M, where M is an integer greater than or equal to 1. The step of obtaining the target fusion data corresponding to the first related image block by fusing the first related image block with the encoding / decoding information corresponding to the first related image block by the information fusion layer of the image filtering processor based on the neural network is as follows: The step of performing a convolution operation on M first related image blocks and encoding / decoding information by an information fusion layer of the image filtering processor based on the neural network to obtain target fusion data corresponding to the first related image blocks is included. The multimedia data processing method according to claim 2.

6. The number of the first related image blocks is M, where M is an integer greater than or equal to 1. The image filtering processor based on the neural network includes M information fusion layers, and one information fusion layer corresponds to one related image block. The step of obtaining target fusion data corresponding to the first related image blocks by fusing the first related image blocks with the encoding / decoding information corresponding to the first related image blocks by an information fusion layer of the image filtering processor based on the neural network is as follows: The step of performing a dot product operation on a first related image block i and the encoding / decoding information corresponding to the first related image block i by an information fusion layer i of the image filtering processor based on the neural network to obtain fusion data corresponding to the first related image block i, where the first related image block i belongs to the M first related image blocks, i is a positive integer less than or equal to M, and the information fusion layer i is the information fusion layer corresponding to the first related image block i among the M information fusion layers. When the fusion data corresponding to each of the M first related image blocks is obtained, the step of performing a convolution operation on the fusion data corresponding to each of the M first related image blocks to obtain the target fusion data corresponding to the first related image blocks is included. The multimedia data processing method according to claim 2.

7. The encoding / decoding information corresponding to the target image block includes at least one of a sequence level quantization parameter, a slice level quantization parameter, a slice level encoding type, a block encoding type, filtering intensity image block information corresponding to the target image block, prediction image block information corresponding to the target image block, and block division image block information of the target image block. The multimedia data processing method according to claim 3.

8. The image block information corresponding to the target image block includes at least one of first color component information, second color component information, and third color component information. The first color component information is used to represent the luminance of the image block, the second color component information and the third color component information are used to represent the chroma of the image block, and the image block information corresponding to the target image block includes any one of filtering intensity image block information, block division image block information, and prediction image block information corresponding to the target image block. The multimedia data processing method according to claim 7.

9. The target encoding / decoding information is used to represent the degree of influence of the first related image block on the target image block. The target encoding / decoding information includes at least one of an influence factor of the first related image block on the target image block, a reference direction of the first related image block with respect to the target image block, block division image block information of the first related image block, a quantization parameter corresponding to the first related image block, reconstruction image block information corresponding to the first related image block, filtering intensity image block information corresponding to the first related image block, and prediction image block information corresponding to the first related image block. The influence factor is determined based on the image distance between the first related image block and the target image block. The multimedia data processing method according to claim 1.

10. The multimedia data processing method When the slice where the target image block is located is an incomplete intra-coded slice, determining the image distance between the first related image block and the target image block based on the image order count corresponding to the first related image block and the image order count corresponding to the target image block; When the slice where the target image block is located is a complete intra-coded slice, determining a predetermined distance value as the image distance between the first related image block and the target image block. The multimedia data processing method according to claim 9.

11. The step of determining the image distance between the first related image block and the target image block based on the image sequence count corresponding to the first related image block and the image sequence count corresponding to the target image block is as follows: Determining a count difference between the image sequence count corresponding to the first related image block and the image sequence count corresponding to the target image block; Determining the absolute value of the count difference as the image distance between the first related image block and the target image block, and includes: The multimedia data processing method according to claim 10.

12. The step of determining the image distance between the first related image block and the target image block based on the image sequence count corresponding to the first related image block and the image sequence count corresponding to the target image block is as follows: Determining a count difference between the image sequence count corresponding to the first related image block and the image sequence count corresponding to the target image block; Calculating a count mapping value corresponding to the absolute value of the count difference, and determining the image distance between the first related image block and the target image block based on the count mapping value, and includes: The multimedia data processing method according to claim 10.

13. The image distance between the first related image block and the target image block is determined based on the time domain layer of the target image block, or the image distance between the first related image block and the target image block is determined based on a level mapping value corresponding to the time domain layer of the target image block. The multimedia data processing method according to claim 9.

14. The image block information corresponding to the first related image block includes at least one of first color component information, second color component information, and third color component information. The first color component information is used to represent the luminance of the image block, the second color component information and the third color component information are used to represent the chroma of the image block, and the image block information corresponding to the first related image block includes any one of the reconstructed image block information, block division image block information, filtering intensity image block information, and predicted image block information corresponding to the first related image block. The multimedia data processing method according to claim 9.

15. The multimedia data processing method is as follows: When the slice where the target image block is located is an incomplete intra-coded slice, determining the reference direction of the first related image block with respect to the target image block based on the relationship between the picture order count corresponding to the first related image block and the picture order count corresponding to the target image block; When the slice where the target image block is located is a complete intra-coded slice, adopting a predetermined direction value to indicate the reference direction of the first related image block with respect to the target image block. The multimedia data processing method according to Claim 9.

16. Based on the first related image block and the target encoding / decoding information, the step of obtaining a filtered image block corresponding to the target image block by performing filtering processing on the target image block by an image filtering processor based on a neural network includes: Obtaining the image distance between the target image block and the first related image block from the target encoding / decoding information; When the image distance is greater than a distance threshold, prohibiting performing filtering processing on the target image block based on the first related image block; When the image distance is less than or equal to the distance threshold, performing filtering processing on the target image block by an image filtering processor based on a neural network based on the first related image block and the target encoding / decoding information, and obtaining a filtered image block corresponding to the target image block. The multimedia data processing method according to Claim 1.

17. Based on the first related image block and the target encoding / decoding information, the step of obtaining a filtered image block corresponding to the target image block by performing filtering processing on the target image block by an image filtering processor based on a neural network includes: Obtaining the image distance between the target image block and the first related image block from the target encoding / decoding information; When the image distance does not belong to the specified image distance, a step of prohibiting performing filtering processing on the target image block based on the first related image block; When the image distance belongs to the specified image distance, based on the first related image block and the target encoding / decoding information, a step of performing filtering processing on the target image block by an image filtering processor based on a neural network to obtain a filtering image block corresponding to the target image block; The multimedia data processing method according to claim 1.

18. The size of the target image block is the same as the size of the first related image block, or the size of the target image block is different from the size of the first related image block; There is the same spatial positional relationship between the target image block and the first related image block. The multimedia data processing method according to claim 1.

19. The size of the target image block is the same as the size of the first related image block, or the size of the target image block is different from the size of the first related image block; There is a different spatial positional relationship between the target image block and the first related image block, and the first related image block is determined based on the motion vector of the target image block. The multimedia data processing method according to claim 1.

20. When the size of the target image block is smaller than the size of the first related image block, based on the first related image block and the target encoding / decoding information, the step of performing filtering processing on the target image block by an image filtering processor based on a neural network to obtain a filtering image block corresponding to the target image block is: A step of dividing the first related image block to obtain N image sub-blocks, where N is a positive integer greater than 1, and the target encoding / decoding information includes encoding / decoding information for representing the influence degree of each of the N image sub-blocks on the target image block; A step of inputting the N image sub-blocks and the encoding / decoding information corresponding to each of the N image sub-blocks into an image filtering processor based on a neural network; Based on the N image sub-blocks and the encoding / decoding information corresponding to each of the N image sub-blocks by the image filtering processor based on the neural network, performing a filtering process on the target image block to obtain a filtered image block corresponding to the target image block; The multimedia data processing method according to claim 18.

21. The multimedia data processing method includes: generating a second related image block for decoding a predicted image block in the multimedia data based on the filtered image block corresponding to the target image block, and storing the second related image block in a decoding cache; or skipping the step of generating a display image block based on the filtered image block corresponding to the target image block and storing the filtered image block corresponding to the target image block in a decoding cache, and transmitting the display image block to a display device, where the display device is configured to display the display image block; The multimedia data processing method according to claim 1.

22. A multimedia data processing apparatus, comprising: a determination module configured to determine a first related image block related to a target image block to be filtered in multimedia data; an acquisition module configured to acquire the first related image block and target encoding / decoding information related to the target image block; a filtering module configured to perform a filtering process on the target image block by an image filtering processor based on the neural network based on the first related image block and the target encoding / decoding information to obtain a filtered image block corresponding to the target image block; A multimedia data processing apparatus.

23. A computer device comprising a processor and a memory, where the processor is connected to the memory, the memory is configured to store program code, and the processor is configured to execute the multimedia data processing method according to any one of claims 1 to 21 by calling the program code.

24. A computer program that causes a processor to execute the multimedia data processing method according to any one of claims 1 to 21.

Citation Information

Patent Citations

  • Encoding program, decoding program, encoding device, decoding device, encoding method, and decoding method

    JP2020198463A

  • FILTERING METHOD, APPARATUS, ENCODER AND COMPUTER STORAGE MEDIUM

    JP2022526107A

  • Multiple Neural Network Models for Filtering During Video Coding

    JP2024501331A

  • Method and apparatus for filtering with mode-aware deep learning

    US20200213587A1

  • Multiple neural network models for filtering during video coding

    US20220215593A1