Multimedia data processing method, apparatus, device, and program
The neural network-based image filtering processor addresses low accuracy in multimedia data processing by utilizing related image blocks and coding/decoding information, improving filtering and encoding quality.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2023-07-07
- Publication Date
- 2026-03-17
AI Technical Summary
Current image filtering methods in multimedia data processing suffer from low accuracy, leading to poor filtering effects and reduced encoding quality.
A multimedia data processing method that utilizes a neural network-based image filtering processor, incorporating multidimensional information such as related image blocks and target coding/decoding information to enhance the filtering process.
Improves the accuracy and quality of image filtering, enhancing the encoding process by leveraging the high-quality advantages of related image blocks.
Smart Images

Figure 0007832367000001 
Figure 0007832367000002 
Figure 0007832367000003
Abstract
Description
[Technical Field]
[0001] (Cross-reference to related applications) This application is filed based on the Chinese patent application filed with the China National Patent Office on September 19, 2022, with application number 202211138569.5, claiming priority from said Chinese patent application, and all contents of said Chinese patent application are incorporated into this application by reference.
[0002] This application relates to the field of multimedia technology, and more particularly to multimedia data processing methods, apparatuses, devices, storage media, and program products thereof. [Background technology]
[0003] In the processing of multimedia data, encoding equipment performs operations such as encoding, transformation, and quantization on the original image to obtain an encoded image. Then, operations such as inverse quantization, inverse transformation, and predictive compensation are performed on the encoded image to obtain a reconstructed image. Compared to the original image, the reconstructed image has some information that differs from the original image due to the effects of quantization, resulting in distortion. Therefore, in order to reduce the degree of distortion in the reconstructed image, it is necessary to perform filtering on the reconstructed image. In practice, the accuracy of current image filtering methods is low, resulting in poor filtering effects. [Overview of the project]
[0004] The embodiments of this application provide a multimedia data processing method, apparatus, device, storage medium, and program product that can improve the accuracy of image filtering and enhance the filtering effect of images.
[0005] In one embodiment of the present application, a multimedia data processing method is provided which is performed by a computer device, and the method is The steps include determining a first related image block related to the target image block to be filtered within the multimedia data, The steps include: obtaining target coding and decoding information related to the aforementioned target image block; The process includes the step of performing a filtering process on the target image block using a neural network-based image filtering processor based on the first related image block and the target coding / decoding information, in order to obtain a filtered image block corresponding to the target image block.
[0006] In one embodiment of the present application, a multimedia data processing device is provided, the device is A decision module configured to determine a first associated image block related to a target image block to be filtered within multimedia data, An acquisition module configured to acquire target coding and decoding information related to the aforementioned target image block, The system includes a filtering module configured to perform filtering on the target image block using a neural network-based image filtering processor based on the first related image block and the target coding / decoding information, thereby obtaining a filtered image block corresponding to the target image block.
[0007] In one embodiment of this application, a computer device is provided comprising: a memory configured to store a computer program; and a processor configured to perform the steps in the above method by calling the computer program.
[0008] In one embodiment of the present application, a computer-readable storage medium is provided which stores a computer program that, when executed by a processor, includes program instructions for causing the processor to perform the steps in the above-described method.
[0009] In one embodiment of the present application, a computer program product is provided which, when executed by a processor, includes a computer program / instruction causing the processor to carry out the steps of the above method.
[0010] The embodiments of this application introduce multidimensional information such as related image blocks and target coding / decoding information to provide a richer amount of information in the filtering process of target image blocks. At the same time, the target coding / decoding information is used to represent the degree of influence of the first related image block on the target image block. Therefore, by using the target coding / decoding information and a neural network-based image filtering processor, filtering is performed on the reconstructed image block, making full use of the high-quality advantages of the related image block. This helps to improve the accuracy and quality of the filtering process on the target image block by using related image blocks selectively in the filtering process on the target image block. In other words, it is possible to improve the accuracy of the filtering process on the image, improve the filtering effect on the image, and improve the coding quality of multimedia data. [Brief explanation of the drawing]
[0011] [Figure 1] This is a flowchart of the video processing described in this application. [Figure 2] This is an illustrative flowchart of the multimedia data processing method according to this application. [Figure 3] This is a schematic diagram of the block division of an image frame according to this application. [Figure 4] This is an illustrative flowchart of the multimedia data processing method described in this application. [Figure 5] This is a schematic diagram illustrating the encoded reference relationship between image frames according to this application. [Figure 6] This is a schematic diagram illustrating the encoded reference relationship between image frames according to this application. [Figure 7]It is a schematic diagram showing the encoding reference relationship between image frames according to the present application. [Figure 8a] It is a schematic diagram of obtaining the first related image block according to the present application. [Figure 8b] It is a schematic diagram of obtaining the first related image block according to the present application. [Figure 8c] It is a schematic diagram of obtaining the first related image block according to the present application. [Figure 9a] It is a schematic diagram of obtaining the first related image block according to the present application. [Figure 9b] It is a schematic diagram of obtaining the first related image block according to the present application. [Figure 9c] It is a schematic diagram of obtaining the first related image block according to the present application. [Figure 10] [[ID=2~1]]It is a schematic diagram of an image filtering processor based on a neural network that performs filtering processing on a target image block according to the present application. [Figure 11] It is a schematic diagram of generating an image filtering processor based on a neural network according to the present application. [Figure 12a] It is an exemplary structural diagram of a Resblock in an image filtering processor based on a neural network according to the present application. [Figure 12b] It is an exemplary structural diagram of a Resblock in an image filtering processor based on a neural network according to the present application. [Figure 13a] It is a schematic diagram of an image filtering processor based on a neural network that performs filtering processing on a target image block according to the present application. [Figure 13b] It is a schematic diagram of an image filtering processor based on a neural network that performs filtering processing on a target image block according to the present application. [Figure 13c] It is a schematic diagram of an image filtering processor based on a neural network that performs filtering processing on a target image block according to the present application. [Figure 13d]This is a schematic diagram of a neural network-based image filtering processor that performs filtering on a target image block, as described in this application. [Figure 13e] This is a schematic diagram of a neural network-based image filtering processor that performs filtering on a target image block, as described in this application. [Figure 13f] This is a schematic diagram of a neural network-based image filtering processor that performs filtering on a target image block, as described in this application. [Figure 13g] This is a schematic diagram of a neural network-based image filtering processor that performs filtering on a target image block, as described in this application. [Figure 14] This is an illustrative structural diagram of the multimedia data processing device described in this application. [Figure 15] This is an illustrative structural diagram of a computer device according to an embodiment of this application. [Modes for carrying out the invention]
[0012] To more clearly illustrate the embodiments of this application or the technical solutions of the prior art, the following is a brief introduction to the drawings used in the description of the embodiments or the prior art. The drawings described below are only a few embodiments of this application, and it will be obvious to those skilled in the art that other drawings can be obtained by following these drawings without any creative work.
[0013] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the drawings of the embodiments of this application, and it is clear that the embodiments described are merely a part of the embodiments of this application, and not all embodiments. All other embodiments obtained by those skilled in the art without creative effort based on the embodiments of this application shall be included within the scope of protection of the present invention.
[0014] The embodiments of this application relate to a multimedia data processing technology. Here, multimedia data (also called media data) refers to composite data formed by media data such as text, graphics, images, audio, video, and dynamic images, where the content is related to each other. The multimedia data referred to in the embodiments of this application mainly includes image data consisting of images, or video data consisting of images and audio. The processing process for multimedia data according to the embodiments of this application mainly includes media data collection, media data encoding, media data file packaging, media data file transmission, media data decoding, and final data display. If the multimedia data is video data, the complete processing process for video data may be as shown in Figure 1, and specifically may include video collection, video encoding, video file packaging, video transmission, video file depackaging, video decoding, and final video display.
[0015] Video acquisition is used to convert analog video to digital video and save it in the form of a digital video file. In other words, video acquisition can convert a video signal into binary digital information, where the binary information converted from the video signal is a binary data stream, which is also called the code stream or bitstream of the video signal. Video coding is the process of converting a file in an original video format to a file in another video format using compression techniques. The generation of video media content referred to in the embodiments of this application includes real scenes generated by camera capture and scenes of screen content generated by a computer. From the perspective of the video signal acquisition method, video signals can be divided into two methods: those captured by a camera and those generated by a computer. Due to differences in statistical characteristics, the corresponding compression encoding methods may also differ. Modern mainstream video encoding technologies employ a hybrid encoding framework, such as the International Video Coding Standard (HEVC: High Efficiency Video Coding) / H.265, the International Video Coding Standard (VVC: Versatile Video Coding) / H.266, and the Chinese National Video Coding Standard (AVS: Audio Video Coding Standard), or AVS3 (the third-generation video coding standard introduced by the AVS Standards Group), and perform the following series of operations and processing on the input original video signal, specifically as shown in Figure 2. (i) Block partition structure: The input multimedia data frame (e.g., one video frame in video data) is divided into several non-overlapping processing units based on the size of one processing unit, and a similar compression operation is performed on each processing unit. In one embodiment, this processing unit is called a Coding Tree Unit (CTU) or Largest Coding Unit (LCU).Further subdivision of the CTU allows for finer divisions, yielding one or more basic coding units called coding units (CUs). Each CU is the most fundamental element in a single coding stage. In another embodiment, this processing unit is also called a coding tile (a rectangular area of a multimedia data frame that can be decoded and coded on its own). Further subdivision of the coding tile allows for finer divisions, yielding one or more superblocks (SBs) (the starting point for block division, which can continue to be divided into multiple subblocks), and then further subdivision of the superblocks yields one or more image blocks (Bs). Each image block is the most fundamental element in a single coding stage.
[0016] Here, the relationship between LCU (or CTU) and CU is as shown in Figure 3, and as can be seen from Figure 3, one image frame in multimedia data contains one or more maximum coding units, and one maximum coding unit contains at least two coding units. Understandably, in the coding process for multimedia data, if the coding device performs block division on the image frames in the multimedia data, one image block in the embodiments of this application may refer to the maximum coding unit (LCU) or coding unit (CU) in the image of one frame of multimedia data. In the coding process for multimedia data, if the coding device does not perform block division on the image frames in the multimedia data, one image block in the embodiments of this application may refer to the image of one frame of multimedia data. (ii) Predictive Coding: Including methods such as intra-prediction and inter-prediction, the original video signal is used to obtain a residual video signal by predicting a selected coded video signal. On the encoding side, it is necessary to select the most suitable predictive coding mode from various possible modes for the currently encoded image block (i.e., the image block to be decoded) and inform the decoding side of it.
[0017] a. Intra (picture) prediction: The predicted signal comes from an encoded and reconstructed region within the same image.
[0018] b. Inter (picture) prediction: The predicted signal comes from another encoded image (called a reference image) that is different from the current image.
[0019] (iii) Transform & Quantization: The residual video signal is transformed into a transformation domain by transformation operations such as the Discrete Fourier Transform (DFT) and the Discrete Cosine Transform (DCT), which are called transformation coefficients. The DCT is a subset of the DFT signal in the transformation domain, and a lossy quantization operation is performed to further reduce the amount of information lost, making the quantized signal more suitable for compressed representation.
[0020] Some video encoding standards may offer multiple transformation options; therefore, the encoding side must select one of these transformations for the currently encoded image block and inform the decoding side. The fineness of quantization is usually determined by the quantization parameter (QP). A large QP value means that coefficients with larger values are quantized to the same output, resulting in greater distortion and a lower code rate. Conversely, a small QP value means that coefficients with smaller values are quantized to the same output, resulting in less distortion and a higher code rate.
[0021] (iv) Entropy Coding or Statistical Coding: The quantized transformation domain signal is statistically compressed and coded based on the frequency of each value, and finally a binarized (0 or 1) compressed code stream is output. Simultaneously, coding generates other information, such as the selected mode and motion vectors, and these also need to be entropy coded to reduce the coding rate.
[0022] Statistical coding is a type of reversible coding scheme that can effectively reduce the coding rate required to represent the same signal. Common statistical coding schemes include Variable Length Coding (VLC) and Content Adaptive Binary Arithmetic Coding (CABAC).
[0023] (v) Loop Filtering: Encoded images (i.e., multimedia data frames) are reconstructed into decoded images by operations of inverse quantization, inverse transform, and predictive compensation (the reverse operations of (ii) to (iv) above). Compared to the original image, the reconstructed image differs from the original image in some information due to the effects of quantization, resulting in distortion. Filtering operations are performed on the reconstructed image, and filters such as deblocking filters, sample adaptive offsets (SAO), or adaptive loop filters (ALF) can be used to effectively reduce the degree of distortion caused by quantization. These filtered reconstructed images are used to predict subsequent signals as a reference for subsequent encoded images, and therefore the above filtering operations are also called loop filtering or filtering operations within the encoding loop.
[0024] Figure 2 shows the basic process of a video encoder, where the kth CU(S k Let's explain using [x,y] as an example, where k is a positive integer greater than or equal to 1 and less than or equal to the number of CUs in the currently input image, S k [x,y] represents the pixel point at coordinates [x,y] in the kth CU, where x is the horizontal coordinate of the pixel point, and y is the vertical coordinate of the pixel point. k [x,y] is predicted by one preferred process, such as motion compensation or intraprediction, resulting in the predicted signal S^ k [x,y] is obtained, S k[x, y] is subtracted from S^ k [x, y] to obtain a residual signal U k [x, y]. Then, the residual signal U k [x, y] is subjected to transformation and quantization. The data output by quantization is output to two different processes. One is sent to an entropy encoder for entropy encoding. The encoded code stream is output to a buffer and saved, waiting to be output. The other is subjected to inverse quantization and inverse transformation to obtain a signal U’ k [x, y]. The signal U’ k [x, y] is added to S^ k [x, y] to obtain a new prediction signal S* k [x, y]. S* k [x, y] is sent to and saved in the buffer of the current image. S* k [x, y] is used to obtain f(S* k [x, y]) by intra-image prediction. S* k [x, y] is loop-filtered to obtain S’ k [x, y]. To generate a reconstructed video, S’ k [x, y] is sent to and saved in the decoded image buffer. S’ k [x, y] is obtained by motion compensation prediction as S’ k r [x + m x,y + m y . S’ r [x + m x,y + m y indicates a reference block, and m x and m y respectively indicate the horizontal and vertical components of the motion vector.
[0025] After encoding multimedia data, the encoded data stream needs to be packaged and transmitted to the user. Video file packaging refers to storing encoded and compressed video and audio in a single file according to a specific format, using a packaging format (or container, or file container). Common packaging formats include Audio Video Interleaved (AVI) or ISO Based Media File Format (ISOBMFF), where ISOBMFF is a media file packaging standard, and the most typical ISOBMFF file is a Moving Picture Experts Group 4 (MP4) file.
[0026] The packaged file is transmitted via video to a decoding device (i.e., a user terminal), where the decoding device performs reverse operations such as depackaging and decoding, and then the decoding device can represent the final video content. Here, the packaged file can be sent to the decoding device via a transmission protocol, which may be Dynamic Adaptive Streaming over HTTP (DASH), a self-adaptive bitrate streaming technology. By transmitting using DASH, high-quality streaming media can be delivered over the internet via conventional HTTP network servers. In DASH, media fragment information is described using Media Presentation Description (MPD) within DASH, and in DASH, one or more combinations of media components, such as a video file of a certain resolution, can be considered as one representation. Multiple representations included are considered as one set of video streams (Adaptation Set), and one DASH may contain one or more Adaptation Sets.
[0027] It can be understood that the depackaging process of a decoding device is the reverse of the file packaging process described above. The decoding device depackages the packaged file according to the file format requirements at the time of packaging, and obtains the audio code stream and video code stream. The decoding process of a decoding device is also the reverse of the encoding process. The decoding device can decode the audio code stream to restore the audio content. As can be seen from the encoding process described above, on the decoding side, for each CU, after the decoder obtains the compressed code stream, it first performs entropy decoding to obtain various mode information and quantized transformation coefficients. For each coefficient, the residual signal is obtained by inverse quantization and inverse transformation. On the other hand, based on known encoding mode information, a prediction signal corresponding to the CU can be obtained, and the two can be added together to obtain the reconstructed signal. Finally, the reconstructed value of the decoded image requires loop filtering to generate the final output signal.
[0028] As can be seen from the above, both the encoding and decoding sides require filtering of the reconstructed image, and because the accuracy of current filtering methods for images is low, the filtering effect on the image is not good. In light of this, the embodiments of this application introduce multidimensional information such as related image blocks and target encoding / decoding information to provide a richer amount of information in the filtering process of the target image block. At the same time, the target encoding / decoding information is used to represent the degree of influence of the first related image block on the target image block. Therefore, by using the target encoding / decoding information to fully utilize the advantages of the high quality of the related image block, filtering is performed on the reconstructed image block, and the related image block is used differentially in the filtering process for the target image block, which helps to improve the accuracy of the filtering process of the target image block and improve the quality of the filtering process of the target image block. In other words, the accuracy of the filtering process of the image can be improved, the filtering effect of the image can be improved, and furthermore, the encoding quality of multimedia data can be improved.
[0029] Furthermore, referring to Figure 4, an exemplary flowchart of a multimedia data processing method according to an embodiment of the present application, the method can be performed by computer equipment, which may refer to encoding equipment or decoding equipment. As shown in Figure 4, the method may include at least the following S101 to S103.
[0030] In S101, a first related image block is determined that is associated with the target image block to be filtered within the multimedia data.
[0031] In S102, target coding and decoding information related to the target image block is obtained.
[0032] In steps S101 and S102, the computer equipment can determine a reconstructed and unfiltered image block in the multimedia data as the target image block to be filtered, or a reconstructed image block that has undergone rudimentary filtering in the multimedia data as the target image block to be filtered, where rudimentary filtering refers to a conventional filtering method. Furthermore, it can acquire a first related image block associated with the target image block, and acquire target coding / decoding information associated with the first related image block and the target image block. The target coding / decoding information includes coding / decoding information corresponding to the first related image block, and the target coding / decoding information may further include coding / decoding information corresponding to the target image block, and the first related image block belongs to a related image that has an coding reference relationship with the target image block, that is, the first related image block belongs to a reference image that has an coding reference relationship with the target image block.
[0033] To understand it, multimedia data contains images of multiple frames, and images of multiple frames are also called sequences of multimedia data. An image of one frame contains one or more slices, and a slice contains one or more image blocks. If an image of one frame contains one slice, and a slice contains one image block, then an image block refers to the image corresponding to that image block. According to the encoding scheme of the slices, slices of an image include fully intra-encoded slices and incompletely intra-encoded slices. A fully intra-encoded slice means that the encoding scheme of all image blocks in the slice is intra-encoded (i.e., intra-predictive), and an incompletely intra-encoded slice means that the encoding scheme of the image blocks in the slice may be inter-encoded (i.e., inter-predictive). If all slices in an image belong to a fully intra-encoded slice, the image is also called an I-frame or keyframe. If an image contains slices that belong to an incompletely intra-encoded slice, the image may be a B-frame or a P-frame. Unencoded image blocks within an image are also called original image blocks; encoded original image blocks within an image are also called encoded image blocks (or predicted image blocks); reconstructed encoded image blocks within an image are also called reconstructed image blocks; unfiltered or partially filtered reconstructed image blocks within an image are also called filtered target image blocks; and filtered target image blocks within an image are also called filtered image blocks.
[0034] It can be understood that in multimedia data encoding technology, time-domain layer partitioning technology is also involved. This technology can partition different image frames into different time-domain layers according to the dependencies during encoding. Specifically, by using this time-domain layering technology to partition the time-domain layers of images in multimedia data, if an image frame is partitioned into a lower layer, it is not necessary to refer to image frames of higher layers during encoding. Here, the time-domain layer of an image can be used to represent the encoding order of the image. For example, the lower the time-domain layer, the earlier the image is encoded, and the higher the time-domain layer, the later the image is encoded. The time-domain layer of an image can also be used to represent the image distance between an image and its reference image. For example, the larger the time-domain layer of an image, the closer the image distance between the image and its reference image, and the smaller the time-domain layer of an image, the greater the image distance between the image and its reference image. All image blocks within the same image have the same time-domain layer, and the time-domain layer of an image block refers to the time-domain layer of the image to which that image block belongs.
[0035] It can be understood that the target image block and the first related image block may belong to the same image in the multimedia data, or they may belong to different images in the multimedia data. For example, if the image to which the target data block belongs is an I-frame, and an I-frame refers to a keyframe, then image frames of this type of media do not need to be encoded depending on other image frames and only the use of an intra encoding scheme is permitted. That is, if the encoding scheme of the target image block is an intra encoding scheme, then the target image block itself can be the first related image block. As shown in Figure 5, the multimedia data contains 9 frames of images, and the encoding scheme of all 9 frames of images is an intra encoding scheme. The encoding process of the image blocks within each frame of images is all dependent on the image blocks within each image. The numbers above the images in each frame indicate the encoding order of the images in each frame. There are no encoding reference relationships between the images in each frame, and the Temporal Layer Id of the images in each frame is all 0.
[0036] In another example, the image to which the target data block belongs is a B-frame or P-frame, also called an inter-encoded frame. Such image frames allow the use of inter-encoded and intra-encoded schemes. In this case, the encoding scheme of the target image block is a non-intra-encoded scheme. In this case, the first related image block belongs to a related image that has an encoding reference relationship with the target image block. As shown in Figure 6, the multimedia data contains 9 frames of images, where 8 frames are B-frames and 1 frame is an I-frame. The numbers above the images in each frame indicate the encoding order of the images in each frame, and the arrows between the images in each frame are used to represent the encoding reference relationships between the images in each frame. In Figure 6, the Temporal Layer Id of all the images in each frame is 0. The media type of the image at the beginning of the encoding order is an I-frame, and the media type of the image at the end of the encoding order is a B-frame. The encoding process for each B-frame all refers to the first I-frame and the corresponding adjacent image frame (i.e., the image frame in which the encoding order is located before it). That is, if the target image block belongs to a B-frame in Figure 6, the first associated image block belongs to the first I-frame or an adjacent image to the image to which the target image block belongs. If the target image block belongs to the image in the third position in the encoding order, there are two first associated image blocks associated with the target image block, for example, first associated image block 1 and first associated image block 2, where first associated image block 1 belongs to the first I-frame and first associated image block 2 belongs to the image in the second position in the encoding order.
[0037] In another example, as shown in Figure 7, the multimedia data contains 9 frames of images, where 8 frames are B-frames and 1 frame is an I-frame. The numbers above each frame's image are used to indicate the encoding order of the image in each frame, and the arrows between each frame's images are used to represent the encoding reference relationships between the images in each frame. In Figure 7, the image at the beginning of the encoding order is an I-frame, and the image at the end of the encoding order is a B-frame. The time-domain layer of the image at the beginning of the encoding order and the image at the second position are both 0, indicating Layer 0, while the time-domain layer of the image at the third position is 1, indicating Layer 1. The time-domain layer of the image at the fourth position and the image at the seventh position are both 2, indicating Layer 2, while the time-domain layer of the image at the fifth, sixth, eighth, and ninth position is all 3, indicating Layer 3. The encoding process for images with a low time-domain layer does not depend on images with a high time-domain layer, while the encoding process for images with a high time-domain layer can depend on images with a low time-domain layer.
[0038] It can be understood that the height of the time-domain layers referred to in the embodiments of this application is a relative concept, as is the case with the four time-domain layers Layer0 to Layer3 determined in Figure 7. For the time-domain layer of Layer0, Layer1 to Layer3 are all high time-domain layers; for the time-domain layer of Layer1, the time-domain layer of Layer3 is a high time-domain layer of Layer1; and the time-domain layer of Layer0 is a low time-domain layer of Layer1. In Figure 7, if the target image block belongs to the image that is first in the coding order, the associated image block related to the target image block also belongs to the image that is first in the coding order. If the target image block belongs to the image that is after the first in the coding order (i.e., the image that is at the beginning of the coding order), the image to which the associated image block related to the target image block belongs is different from the image to which the target image block belongs. If the target image block belongs to the image that is last in the coding order, the associated image block related to the target image belongs to the image that is first in the coding order.
[0039] In essence, the target encoding / decoding information is used to represent the degree of influence of the first related image block on the target image block. A higher degree of influence of the first related image block on the target image block, as expressed in the target encoding / decoding information, indicates that the encoding process of the target image block will refer to more information from the first related image block. Therefore, the filtering process applied to the target image block will also refer to more information from the first related image block. Conversely, a lower degree of influence of the first related image block on the target image block, as expressed in the target encoding / decoding information, indicates that the encoding process of the target image block will refer to less information from the first related image block. Therefore, the filtering process applied to the target image block will also refer to less information from the first related image block. The target decoding information allows for the differentiated use of related image blocks, thereby improving the accuracy of the filtering process applied to the target image block.
[0040] It can be understood that the target coding / decoding information includes coding / decoding information corresponding to the first related image block, or the target coding / decoding information includes coding / decoding information corresponding to the first related image block and coding / decoding information corresponding to the target image block. The coding / decoding information corresponding to the target image block includes at least one of the following: sequence-level quantization parameters, slice-level quantization parameters, slice-level coding type, block coding type, filtering intensity image block information of the target image block, predicted image block information corresponding to the target image block, and block division image block information of the target image block. Here, the predicted image block information corresponding to the target image block includes all or part of the information of the predicted image block corresponding to the target image block. For example, the predicted image block information corresponding to the target image block includes at least one of the chroma component information and luminance component information of the predicted image block corresponding to the target image block, and the predicted image block corresponding to the target image block refers to the one obtained by performing predictive coding on the original image block corresponding to the target image block. The filtering intensity image block information of the target image block refers to at least one of the chroma component information and luminance component information of the filtering intensity image block information corresponding to the target image block. The block division image block information of the target image block includes all or part of the information of the image blocks in the image generated by block division. For example, the block division image block information of the target image block includes at least one of the chroma component information and luminance component information of the image blocks in the image generated by block division. The filtering intensity image block here is an image block generated based on the filtering intensity determination result of the deblocking filter in the codec. The block division image block information is an image block generated based on the encoded block division result. A sequence refers to a series of encoded frames, and a slice is a multimedia data encoding process in which an image is divided into multiple slices, similar to a tile. If not divided, one frame may be considered as one slice.
[0041] It can be understood that the encoding and decoding information corresponding to a target image block has some influence on the image quality of that block. Therefore, the encoding and decoding information corresponding to a target image block is used as auxiliary information for the filtering process of the target image block. The more information contained in the encoding and decoding information corresponding to a target image block, the more it helps to assist in the filtering process of the target image block and improve the accuracy of the filtering process.
[0042] It can be understood that the encoded / decoded information corresponding to the first related image block includes at least one of the following: an influence factor of the first related image block on the target image block, a reference direction of the first related image block on the target image block, block division image block information of the first related image block, quantization parameters corresponding to the first related image block, reconstructed image block information corresponding to the first related image block, filtering intensity image block information corresponding to the first related image block, and predicted image block information corresponding to the first related image block. The encoded / decoded information of the first related image block is used to represent the degree of influence of the first related image block on the target image block. For example, the influence factor is determined based on the image distance between the first related image block and the target image block, and the influence factor is used to represent the degree of influence of the first related image block on the target image block. A larger influence factor indicates a higher degree of influence of the first related image block on the target image block, and a smaller influence factor indicates a lower degree of influence of the first related image block on the target image block. The influencing factor is determined based on the image distance between the first related image block and the target image block. There is an inverse correlation between the influencing factor and the image distance; that is, the greater the image distance between the first related image block and the target image block, the smaller the influencing factor of the first related image block on the target image block. In other words, the closer the image distance between the first related image block and the target image block, the larger the influencing factor of the first related image block on the target image block.
[0043] In another example, the quantization parameter corresponding to the first related image block is used to represent the degree of influence of the first related image block on the target image block. A larger quantization parameter corresponds to the first related image block, the less information the first related image block itself contains, and the lower the degree of influence of the first related image block on the target image block. A smaller quantization parameter corresponds to the first related image block, the more information the first related image block itself contains, and the higher the degree of influence of the first related image block on the target image block. In another example, the image block information corresponding to the first related image block is used to represent the degree of influence of the first related image block on the target image block. A larger amount of information in the image block information corresponding to the first related image block results in a higher degree of influence of the first related image block on the target image block, while a smaller amount of information in the image block information corresponding to the first related image block results in a lower degree of influence of the first related image block on the target image block. The image block information corresponding to the first related image block includes at least one of the following: reconstructed image block information, block-divided image block information, filtering intensity image block information, and predicted image block information. It can be understood that the reference direction of the first related image block to the target image block can be used as one piece of auxiliary information for the target image block filtering process, providing more auxiliary information for the target image block filtering process and improving the filtering effect.Understandably, if the slice in which the target image block is located is an incomplete intra-encoded slice, the computer equipment can determine the image distance between the first related image block and the target image block based on the image order count corresponding to the first related image block and the image order count corresponding to the target image block, where the picture order count (POC) is used to represent the display order of the image to which the image block belongs, and if the slice in which the target image block is located is a complete intra-encoded slice, a predetermined distance value is determined as the image distance between the first related image block and the target image block, where the predetermined distance value may refer to 0, 1025, or other specific values. Here, if the slice in which the target image block is located is a fully intra-coded slice, the slice in which the target image block is located may be called an I-slice, and if the slice in which the target image block is located is an incomplete intra-coded slice, the slice in which the target image block is located may be called a P-slice or a B-slice. By generating the image distance between the target image block and the associated image block using the above method, the I-slice, B-slice, and P-slice can share the above-mentioned neural network-based image filtering processor to perform filtering, saving resources and improving the versatility of the neural network-based image filtering processor without having to train a separate neural network-based image filtering processor for each of the I-slice, B-slice, and P-slice.
[0044] It can be understood that the image sequence count corresponding to the target image block refers to the image sequence count of the image frame to which the target image block belongs, and the image sequence count corresponding to the first related image block refers to the image sequence count of the image frame to which the first related image block belongs. It can be understood that a computer device can determine the image distance between the first related image block and the target image block based on the image sequence count corresponding to the first related image block and the image sequence count corresponding to the target image block, via one of the following four methods.
[0045] In method 1, the computer device can calculate the count difference between the image sequence count corresponding to the first related image block and the image sequence count corresponding to the target image block, and determine the absolute value of this count difference as the image distance between the first related image block and the target image block.
[0046] In method 2, the computer device obtains the count difference between the image sequence count corresponding to the first related image block and the image sequence count corresponding to the target image block, calculates a count mapping value corresponding to the absolute value of the count difference, and determines the image distance between the first related image block and the target image block based on the count mapping value. The count mapping here may be (N - image distance) / M, where M may be a number other than zero, and N may be any value, for example, N is 1025 and M is 1024.
[0047] In method 3, the computer equipment can acquire the time-domain layer of the target image block and determine the image distance between the first related image block and the target image block based on the time-domain layer of the target image block. That is, the image distance between the first related image block and the target image block is inversely correlated with the time-domain layer of the target image block; the higher the time-domain layer of the target image block, the closer the image distance between the first related image block and the target image block becomes, and conversely, the lower the time-domain layer of the target image block, the farther the image distance between the first related image block and the target image block becomes. For example, the difference between N and the time-domain layer of the target image block can be determined as the image distance between the first related image block and the target image block, and N minus the time-domain layer of the target image block can be determined as the image distance between the first related image block and the target image block.
[0048] In method 4, the computer equipment can determine a level mapping value corresponding to the time-domain layer of the target image block and determine the image distance between the first related image block and the target image block based on the level mapping value. For example, the ratio of the time-domain layer of the target image block to a predetermined time-domain layer value is determined as the level mapping value corresponding to the time-domain layer of the target image block, for example, the predetermined time-domain layer value is K and is a non-zero number. Alternatively, the difference between the predetermined time-domain layer value and the time-domain layer of the target image block is determined as the level mapping value corresponding to the time-domain layer of the target image block.
[0049] As can be understood, the image block information corresponding to the first related image block includes at least one of the first color component information, the second color component information, and the third color component information, wherein the first color component information is used to represent the brightness of the image block, and the second and third color component information are used to represent the chroma of the image block, and the image block information corresponding to the first related image block includes one of the reconstructed image block information, block division image block information, filtering intensity image block information, and predicted image block information corresponding to the first related image block.
[0050] In other words, if the slice in which the target image block is located is an incomplete intra-coded slice, the computer device determines the reference direction of the first related image block to the target image block based on the relationship between the magnitude of the image sequence count corresponding to the first related image block and the image sequence count corresponding to the target image block. If the slice in which the target image block is located is a complete intra-coded slice, a predetermined direction value is adopted to indicate the reference direction of the first related image block to the target image block. In other words, by determining the reference direction of the first related image block to the target image block using the above method, I-slice, B-slice, and P-slice can share the above-mentioned neural network-based image filtering processor to perform filtering, saving resources and improving the versatility of the neural network-based image filtering processor without having to train a separate neural network-based image filtering processor for each of the I-slice, B-slice, and P-slice.
[0051] For example, if the slice in which the target image block is located is an incomplete intra-encoded slice and the image sequence count corresponding to the first related image block is greater than the image sequence count corresponding to the target image block, a first direction value is adopted to indicate the reference direction of the first related image block to the target image block. If the slice in which the target image block is located is an incomplete intra-encoded slice and the image sequence count corresponding to the first related image block is less than the image sequence count corresponding to the target image block, a second direction value is adopted to indicate the reference direction of the first related image block to the target image block. The first direction value may be greater than zero, and the second direction value may be less than zero. Unlike the second direction value, the first direction value may be 1, and the second direction value may be 0. If the slice in which the target image block is located is a fully intra-encoded slice, a predetermined direction value is adopted to indicate the reference direction of the first related image block to the target image block. The predetermined direction value may be 0, 1, 2, or -1, and may be the same as the first or second direction value, or may be different from either the first or second direction value.
[0052] It can be understood that the selection method for the first associated image block of the target image block may refer to selecting one or more image blocks as the first associated image block from a specified reference direction. If there are insufficient first associated image blocks, they are filled with reusable first associated image blocks. For example, the selection method for the first associated image block may include three methods such as 1) selecting one or two in the L0 direction, 2) selecting one or two in the L1 direction, and 3) selecting one or two in the L0 direction and one or two in the L1 direction, where the first associated image block includes the image block in the first reference frame of the reference frame list in the L0 direction and the image block in the first reference frame of the reference frame list in the L1 direction.
[0053] Here, L0 may point in the direction where the image order count (POC) of the image frame to which the first related image block belongs is smaller than the image order count (POC) of the image frame to which the target image block belongs, and L1 may point in the direction where the image order count (POC) of the image frame to which the first related image block belongs is larger than the image order count (POC) of the image frame to which the target image block belongs. Alternatively, both L0 and L1 may point in the direction where the image order count (POC) of the image frame to which the first related image block belongs is smaller than the image order count (POC) of the image frame to which the target image block belongs.
[0054] It can be understood that, when the target image block and the first related image block have the same spatial positional relationship, the size of the target image block is either the same as the size of the first related image block, or the size of the target image block is different from the size of the first related image block. For example, the image to which the target image block belongs is the target image, and the image to which the first related image block belongs is the reference image. In Figures 8a, 8b, and 8c, all gray areas in the target image are target image blocks, the gray areas in the reference image of Figure 8a are first related image blocks, and the gray areas and diagonally filled areas in the reference images of Figures 8b and 8c are all first related image blocks. In Figures 8a, 8b, and 8c, the target image block and the first related image block have the same spatial positional relationship. Here, the size of the target image block in Figure 8a is the same as the size of the first related image block, and the size of the target image block in Figures 8b and 8c is smaller than the size of the first related image block.
[0055] It can be understood that there is a different spatial relationship between the target image block and the first related image block, and if the first related image block is determined based on the motion vector of the target image block, then the size of the target image block is either the same as the size of the first related image block, or the size of the target image block is different from the size of the first related image block. The motion vector of the target image block refers to one determined by the motion vectors of each encoded block within the target image block, such as the average motion vector.
[0056] For example, the image to which the target image block belongs is the target image, the image to which the first related image block belongs is the reference image, the direction indicated by the arrows in Figures 9a, 9b, and 9c is the direction of the motion vector of the target image block, and the first related image block is obtained by selecting it from the reference image based on the position indicated by the motion vector of the target image block. All gray areas in the target image in Figures 9a, 9b, and 9c are target image blocks, the gray areas in the reference image in Figure 9a are first related image blocks, and all gray areas and diagonally filled areas in the reference images in Figures 9b and 9c are first related image blocks. In Figures 9a, 9b, and 9c, there is a different spatial relationship between the target image block and the first related image block, and the first related image block is located in the reference image at the position indicated by the motion vector of the target image block. Here, the size of the target image block in Figure 9a is the same as the size of the first related image block, while the size of the target image block in Figures 9b and 9c is smaller than the size of the first related image block.
[0057] In S103, based on the first related image block and the target coding / decoding information, a neural network-based image filtering processor performs filtering on the target image block to obtain a filtered image block corresponding to the target image block.
[0058] In the embodiments of this application, the computer device can perform filtering on the target image block using a neural network-based image filtering processor based on the first related image block and the target coding / decoding information, thereby obtaining a filtered image block corresponding to the target image block. For example, if the degree of influence of the first related image block on the target image block as expressed in the target coding / decoding information is high, the computer device can assign a larger influence factor (i.e., a weight) to the first related image block in the filtering process for the target image block, thereby helping to increase the importance of the first related image block in the filtering process for the target image block. Conversely, the lower the degree of influence of the first related image block on the target image block as expressed in the target coding / decoding information, the lower the influence of the first related image block on the target image block as expressed in the target coding / decoding information, the computer device can assign a smaller influence factor (i.e., a weight) to the first related image block in the filtering process for the target image block, thereby helping to weaken the importance of the first related image block in the filtering process for the target image block. The image filtering processor, based on target coding and decoding information and a neural network, fully utilizes the high-quality advantages of the associated image blocks to perform filtering on the reconstructed image blocks. In the filtering process for the target image blocks, the associated image blocks are used differentially, which helps to improve the accuracy of filtering the target image blocks.
[0059] In essence, the computer device uses the neural network-based image filtering processor described above to perform filtering on the target image block based on the first related image block and the target coding / decoding information, thereby obtaining a filtered image block corresponding to the target image block. Here, the neural network-based image filtering processor may also refer to an in-loop filter for performing filtering on an image block. For example, the neural network-based image filtering processor may be a neural network-based in-loop filter (NNLF). As shown in Figure 10, the computer device can input the target image block to be filtered, the first related image block, and the target coding / decoding information into the neural network-based image filtering processor. The neural network-based image filtering processor then performs filtering on the target image block based on the first related image block and the target coding / decoding information, and outputs a filtered image block. The filtered image block is the filtered image block corresponding to the target image block.
[0060] As can be understood, as shown in Figure 11, the process of generating the neural network-based image filtering processor involves the following three steps: 1) Encoding the training sequence using encoding software to generate a training set. 2) Training an initial neural network-based image filtering processor based on the training set to obtain the neural network-based image filtering processor. 3) Integrating the neural network-based image filtering processor into software, the software into which the neural network-based image filtering processor is integrated performs filtering on the target image block. The training process of the initial neural network-based image filtering processor uses a loss function, which measures the difference between the predicted value (filtered image corresponding to the sample image block) and the true value (original image block corresponding to the sample image block). The larger the loss value, the larger the difference, and the purpose of training is to reduce the loss. In the case of deep learning-based encoding tools, commonly used loss functions are the L1 norm loss function, the L2 norm loss function, and the smooth L1 loss function.
[0061] It can be understood that a neural network-based image filtering processor may include multiple convolutional layers, multiple activation layers, multiple residual (resblock) layers, and one data shuffle (shuffle) layer, where the convolutional layers are used to extract different input features, the activation layers perform a non-linear mapping on the output of the convolutional layers, the Resblock layers are the basic modules of the neural network-based image filtering processor, and the shuffle is used to obtain data that matches the information dimension of the original image block corresponding to the target image block through adjustment.
[0062] One thing that can be understood is that the number of channels in each convolutional layer (convolutional module) is exactly the same; that is, the number of input and output channels in all convolutional layers of the Resblock layer is the same. In Figure 12a, the Resblock layer contains two 3x3 convolutional layers, and the number of input and output channels in both 3x3 convolutional layers is exactly the same; for example, the number of input and output channels are 64 and 64, respectively.
[0063] One thing that can be understood is that the number of channels in each convolutional layer can be different, meaning that the number of input and output channels in all the convolutional layers of the Resblock layer can be different. In Figure 12b, the Resblock layer contains two 1x1 convolutional layers and one 3x3 convolutional layer, and the number of input and output channels in the two 1x1 convolutional layers and the 3x3 convolutional layer are different. For example, the number of channels can be increased before the activation layer and decreased after the activation layer, with the first 1x1 convolutional layers having 64 and 160 input and output channels, the second 1x1 convolutional layers having 160 and 64 input and output channels, and the 3x3 convolutional layer having 64 and 64 input and output channels, respectively.
[0064] It can be understood that a computer device can perform filtering on a target image block using one of the following two methods:
[0065] In Method 1, the target encoding / decoding information includes encoding / decoding information corresponding to the first related image block, and the encoding / decoding information corresponding to the first related image block refers to parameters correlated with the encoding / decoding process of the first related image block. The computer equipment can obtain target fused data corresponding to the first related image block by fusing the first related image block with the encoding / decoding information corresponding to the first related image block using the information fusion layer of the neural network-based image filtering processor. Furthermore, the computer equipment can perform filtering on the target image block based on the target fused data corresponding to the first related image block using the filtering processing layer of the neural network-based image filtering processor, thereby obtaining a filtered image block corresponding to the target image block. The encoding / decoding information corresponding to the first related image block helps to fully utilize the advantages of the high quality of the first related image block and improve the accuracy and quality of the filtering process on the target image block.
[0066] In method 2, the target encoding / decoding information includes encoding / decoding information corresponding to the first related image block and encoding / decoding information corresponding to the target image block. Specifically, the computer equipment can obtain target fused data corresponding to the first related image block by fusing the first related image block with the encoding / decoding information corresponding to the first related image block using the information fusion layer of the neural network-based image filtering processor. Furthermore, the computer equipment can perform filtering on the target image block using the filtering processing layer of the neural network-based image filtering processor based on the target fused data corresponding to the first related image block and the encoding / decoding information corresponding to the target image block, thereby obtaining a filtered image block corresponding to the target image block. The encoding / decoding information corresponding to the first related image block and the encoding / decoding information corresponding to the target image block fully utilize the advantages of the high quality of the first related image block, provide more information to the filtering process of the target image block, and help improve the accuracy and quality of the filtering process on the target image block.
[0067] It can be understood that the information fusion layer of a neural network-based image filtering processor may refer to the convolutional layer of the neural network-based image filtering processor, or it may refer to the dot product module and convolutional layer of the neural network-based image filtering processor.
[0068] It can be understood that the number of the above-mentioned first related image blocks is M (where M is an integer greater than or equal to 1), and the computer equipment can obtain target fused data corresponding to the first related image blocks by fusing these first related image blocks with the corresponding encoded / decoded information using one of the following three fusing methods.
[0069] In fusion method 1, the neural network-based image filtering processor includes M information fusion layers, each corresponding to one associated image block. When there is one first associated image block, the neural network-based image filtering processor includes one information fusion layer, and the computer device can perform a convolution operation on the first associated image block and the encoded / decoded information corresponding to the first associated image block using the information fusion layer of the neural network-based image filtering processor to obtain target fused data corresponding to the first associated image block. When there are at least two first associated image blocks, the computer device can perform a convolution operation on the first associated image block 1 and the encoded / decoded information corresponding to the first associated image block 1 using the information fusion layer 1 of the neural network-based image filtering processor to obtain fused data corresponding to the first associated image block 1. The information fusion layer 1 of the neural network-based image filtering processor is the information fusion layer corresponding to the first associated image block 1 among the M information fusion layers. The information fusion layer 2 of the neural network-based image filtering processor performs a convolution operation on the first related image block 2 and the encoded / decoded information corresponding to the first related image block 2 to obtain fused data corresponding to the first related image block 2. The information fusion layer 2 of the neural network-based image filtering processor is the information fusion layer corresponding to the first related image block 2 among M information fusion layers. The above steps are repeated until fused data corresponding to each of the M first related image blocks is obtained. When fused data corresponding to each of the M first related image blocks is obtained, the fused data corresponding to each of the M first related image blocks is determined as the target fused data corresponding to the first related image block, and the encoded / decoded information of each first related image block is fused with the first related image block using a convolution operation method, which helps to improve the accuracy and quality of the filtering process on the target image block.
[0070] For example, as shown in Figure 13a, a neural network-based image filtering processor includes at least two convolutional layers, N Resblock layers, multiple activation layers, and one shuffle layer, where the at least two convolutional layers are denoted as convolutional layer 0, convolutional layer 1, convolutional layer 2, convolutional layer 3, ..., convolutional layer N, and the activation layers include activation layer 0, activation layer 1, activation layer 2, .... If the target coding / decoding information includes coding / decoding information corresponding to the first related image block, the first related image block includes related image block 0, and the coding / decoding information corresponding to related image block 0 is coding / decoding information 0. The computer device inputs related image block 0 and coding / decoding information 0 into convolutional layer 0, performs a convolution operation on related image block 0 and coding / decoding information 0 using convolutional layer 0, obtains fused data corresponding to related image block 0, and can determine the fused data corresponding to related image block 0 as the target fused data corresponding to the first related image block. Convolutional layer 1 performs a convolution operation on the target image block, and convolutional layer 2 performs a convolution operation on the processing result for the target image block output by activation layer 1 and the processing result for the target fused data output by activation layer 0.
[0071] For example, as shown in Figure 13b, a neural network-based image filtering processor includes at least two convolutional layers, N Resblock layers, and a plurality of activation layers and one shuffle layer, where the at least two convolutional layers are denoted as convolutional layer 0, convolutional layer 1, convolutional layer 2, convolutional layer 3, ..., convolutional layer N, and the activation layers include activation layer 0, activation layer 1, activation layer 2, .... If the target coding / decoding information includes coding / decoding information corresponding to the first related image block and coding / decoding information corresponding to the target image block, the first related image block includes related image block 0 and related image block 1, the coding / decoding information corresponding to related image block 0 is coding / decoding information 0, and the coding / decoding information corresponding to related image block 1 is coding / decoding information 1. Convolutional layer 0 corresponds to related image block 0, and convolutional layer 1 corresponds to related image block 1. The computer equipment inputs related image block 0 and encoding / decoding information 0 into convolution layer 0, and convolution layer 0 performs a convolution operation on related image block 0 and encoding / decoding information 0 to obtain fused data corresponding to related image block 0. It then inputs related image block 1 and encoding / decoding information 1 into convolution layer 1, and convolution layer 1 performs a convolution operation on related image block 1 and encoding / decoding information 1 to obtain fused data corresponding to related image block 1. The fused data corresponding to related image block 0 and the fused data corresponding to related image block 1 are determined as target fused data corresponding to the first related image block. In Figure 13b, the encoding / decoding information corresponding to the target image block includes sequence-level quantization parameters, slice-level quantization parameters, slice-level encoding type, block encoding type, and a predicted image block corresponding to the target image block.
[0072] In fusion method 2, the neural network-based image filtering processor includes one information fusion layer. The computer device performs a convolution operation on M first related image blocks and encoded / decoded information using the information fusion layer of the neural network-based image filtering processor to obtain target fused data corresponding to the first related image blocks. Specifically, M first related image blocks and encoded / decoded information are simultaneously input to the information fusion layer of the neural network-based image filtering processor, and the information fusion layer of the neural network-based image filtering processor performs a convolution operation on the M first related image blocks and encoded / decoded information to obtain target fused data corresponding to the first related image blocks. The convolution operation method fuses the encoded / decoded information of each first related image block with the first related image block, which helps to improve the accuracy and quality of the filtering process on the target image block.
[0073] For example, as shown in Figure 13c, a neural network-based image filtering processor includes at least two convolutional layers, N Resblock layers, and a plurality of activation layers and one shuffle layer, where the at least two convolutional layers are denoted as convolutional layer 0, convolutional layer 1, convolutional layer 2, convolutional layer 3, ..., convolutional layer N, and the activation layers include activation layer 0, activation layer 1, activation layer 2, .... If the target coding / decoding information includes coding / decoding information corresponding to the first related image block and coding / decoding information corresponding to the target image block, the first related image block includes related image block 0 and related image block 1, the coding / decoding information corresponding to related image block 0 is coding / decoding information 0, and the coding / decoding information corresponding to related image block 1 is coding / decoding information 1. The computer equipment inputs all of the related image block 0, related image block 1, coding / decoding information 0, and coding / decoding information 1 into the convolutional layer 0. The convolutional layer 0 performs a convolution operation on related image block 0, related image block 1, coding / decoding information 0, and coding / decoding information 1 to obtain target fused data corresponding to the first related image block. In Figure 13c, the coding / decoding information corresponding to the target image block includes sequence-level quantization parameters, slice-level quantization parameters, slice-level coding type, block coding type, and a predicted image block corresponding to the target image block.
[0074] In fusion method 3, the neural network-based image filtering processor includes M information fusion layers, each corresponding to one associated image block. When there is one first associated image block, the neural network-based image filtering processor includes one information fusion layer. The computer device uses the information fusion layer of the neural network-based image filtering processor to perform a dot product operation on the first associated image block and the encoded / decoded information corresponding to the first associated image block, thereby obtaining target fused data corresponding to the first associated image block. When there are at least two first associated image blocks, the computer device uses the information fusion layer 1 of the neural network-based image filtering processor to perform a dot product operation on the first associated image block 1 and the encoded / decoded information corresponding to the first associated image block 1, thereby obtaining fused data corresponding to the first associated image block 1. The information fusion layer 1 of the neural network-based image filtering processor is the information fusion layer corresponding to the first associated image block 1 among the M information fusion layers. The information fusion layer 2 of the neural network-based image filtering processor performs a dot product operation on the first related image block 2 and the encoded / decoded information corresponding to the first related image block 2 to obtain fused data corresponding to the first related image block 2. The information fusion layer 2 of the neural network-based image filtering processor is the information fusion layer corresponding to the first related image block 2 among M information fusion layers. The above steps are repeated until fused data corresponding to each of the M first related image blocks is obtained. Once fused data corresponding to each of the M first related image blocks is obtained, a convolution operation is performed on the fused data corresponding to each of the M first related image blocks to obtain target fused data corresponding to the first related image block. Fusing the encoded / decoded information of each first related image block with the first related image block using the dot product operation method helps to improve the accuracy and quality of the filtering process on the target image block.
[0075] For example, as shown in Figure 13d, the neural network-based image filtering processor includes at least two convolutional layers, N Resblock layers, a plurality of activation layers, and one shuffle layer, where the at least two convolutional layers are denoted as convolutional layer 0, convolutional layer 1, convolutional layer 2, convolutional layer 3, ..., convolutional layer N, and the activation layers include activation layer 0, activation layer 1, activation layer 2, ... If the target coding / decoding information includes coding / decoding information corresponding to the first related image block, the first related image block includes related image block 0, and the coding / decoding information corresponding to related image block 0 is coding / decoding information 0. The computer equipment inputs the associated image block 0 and the encoding / decoding information 0 into a dot product module, which performs a dot product operation on the associated image block 0 and the encoding / decoding information 0 to obtain fused data corresponding to the associated image block 0, which is then input into a convolution layer 0, which performs a convolution operation on the fused data corresponding to the associated image block 0 to obtain target fused data corresponding to the first associated image block.
[0076] As can be understood, step S102 above includes the following steps: The computer device obtains the image distance between the target image block and the first related image block from the target coding / decoding information. If the image distance is greater than the distance threshold, it indicates that the similarity between the target image block and the first related image block is relatively low, i.e., that the significance of the target image block referring to the first related image block is low. In this case, filtering processing is prohibited for the target image block based on the first related image block. If the image distance is less than or equal to the distance threshold, it indicates that the similarity between the target image block and the first related image block is relatively high, i.e., that the significance of the target image block referring to the first related image block is high. Therefore, the computer device can perform filtering processing on the target image block based on the first related image block and obtain a filtered image block corresponding to the target image block. By performing filtering processing on target image blocks whose image distance is less than or equal to the distance threshold, it is possible to improve the filtering performance of image blocks by utilizing the first related image block in a targeted manner.
[0077] It is understood that when determining the image distance between a first related image block and a target image block based on the image sequence count of the first related image block and the image sequence count of the target image block, if the distance threshold is 2 and the absolute value corresponding to the difference between the image sequence count of the first related image block and the image sequence count of the target image block is 2 or less, then a filtering process is performed on the target image block based on the first related image block and the target encoding / decoding information to obtain a filtered image block corresponding to the target image block. If the absolute value corresponding to the difference between the image sequence count of the first related image block and the image sequence count of the target image block is greater than 2, then it is prohibited to perform a filtering process on the target image block based on the first related image block to obtain a filtered image block corresponding to the target image block.
[0078] It can be understood that, if the image distance between the first related image block and the target image block is determined based on the time-domain layer of the target image block, then the image distance between the first related image block and the target image block is inversely correlated with the time-domain layer of the target image block. For example, if the distance threshold is 4 and the time-domain layer of the target image block is 4 or greater, then a filtering process is performed on the target image block based on the first related image block and the target coding / decoding information to obtain a filtered image block corresponding to the target image block. If the time-domain layer of the target image block is less than 4, then it is prohibited to perform a filtering process on the target image block based on the first related image block to obtain a filtered image block corresponding to the target image block.
[0079] For example, as shown in Figure 13e, the neural network-based image filtering processor includes at least two convolutional layers, N Resblock layers, and a plurality of activation layers and one shuffle layer, where the at least two convolutional layers are denoted as convolutional layer 0, convolutional layer 1, convolutional layer 2, convolutional layer 3, ..., convolutional layer N, and the activation layers include activation layer 0, activation layer 1, activation layer 2, ... If the image distance between the first related image block and the target image block is greater than the distance threshold, the computer device can prohibit filtering the target image block based on the first related image block. If the image distance between the first related image block and the target image block is less than or equal to the distance threshold, the computer device can input the target coding / decoding information and the first related image block to the neural network-based image filtering processor, where the target coding / decoding information includes coding / decoding information corresponding to the target image block, and where the first related image includes related image block 0 and related image block 1. Specifically, the convolutional layer 0 of the neural network-based image filtering processor performs a convolution operation on related image 0 to obtain fused data corresponding to related image block 0, and the convolutional layer 1 of the neural network-based image filtering processor performs a convolution operation on related image 1 to obtain fused data corresponding to related image block 1. The fused data corresponding to related image block 0 and the fused data corresponding to related image block 1 are determined as the target fused data corresponding to the related image block. In Figure 13e, the encoding and decoding information corresponding to the target image block includes sequence-level quantization parameters, slice-level quantization parameters, slice-level encoding type, block encoding type, and the predicted image block corresponding to the target image block.
[0080] In another example, as shown in Figure 13f, the neural network-based image filtering processor includes at least two convolutional layers, N Resblock layers, and a plurality of activation layers and one shuffle layer, where the at least two convolutional layers are denoted as convolutional layer 0, convolutional layer 1, convolutional layer 2, convolutional layer 3, ..., convolutional layer N, and the activation layers include activation layer 0, activation layer 1, activation layer 2, ... If the image distance between the first related image block and the target image block is greater than the distance threshold, the computer device may prohibit filtering the target image block based on the first related image block. If the image distance between the first related image block and the target image block is less than or equal to the distance threshold, the computer device may input the target coding / decoding information and the first related image block to the neural network-based image filtering processor, where the target coding / decoding information includes coding / decoding information corresponding to the target image block, and where the first related image includes related image block 0 and related image block 1. Specifically, the convolutional layer 0 of the neural network-based image filtering processor performs convolution operations on related image 0 and related image 1 to obtain target fused data corresponding to the related image block. In Figure 13f, the coding and decoding information corresponding to the target image block includes sequence-level quantization parameters, slice-level quantization parameters, slice-level coding type, block coding type, and a predicted image block corresponding to the target image block.
[0081] As can be understood, step S102 above includes the following steps: The computer equipment can obtain the image distance between the target image block and the first related image block from the target coding / decoding information, and if the image distance does not belong to a specified image distance, it prohibits filtering the target image block based on the first related image block, and if the image distance belongs to a specified image distance, it performs filtering on the target image block based on the first related image block and the target coding / decoding information to obtain a filtered image block corresponding to the target image block. For example, the specified image distance may refer to the time domain layer of the target image block being a specified time domain layer, for example, the specified time domain layer may be a layer that is a multiple of 2, Layer 5, Layer 4 and Layer 5, or the specified image distance may refer to the difference between the image sequence count of the target image block and the image sequence count of the first related image block being a specified count difference, for example, the specified count difference may be 1, or 1 and 2, etc. By using the image distance between the first related image block and the target image block, it is possible to narrow down the target and utilize the first related image block to improve the filtering performance of the image block.
[0082] It can be understood that if the size of the target image block is smaller than the size of the first related image block, step S102 above includes the following steps: The computer equipment can divide the first related image block to obtain N image subblocks (where N is a positive integer greater than 1), and the target coding / decoding information includes coding / decoding information to represent the degree of influence of each of the N image subblocks on the target image block. The sizes of the N image subblocks may be the same or different. For example, as shown in Figure 8b, the computer equipment can divide the first related image block into 5 image subblocks, where the 5 image subblocks include the gray area and 4 diagonal-filled areas in the reference image, and the 5 image subblocks are the same size. Furthermore, the N image subblocks and the coding / decoding information corresponding to each of the N image subblocks are input to a neural network-based image filtering processor. The computer device, using an image filtering processor based on the neural network, can perform filtering on the target image block based on the N image subblocks and the encoding / decoding information corresponding to each of the N image subblocks, thereby obtaining a filtered image block corresponding to the target image block. By introducing related image blocks from a wider area, more information is provided to the encoding process of the target image block, improving the accuracy of the filtering process on the target image.
[0083] One possible explanation is that, after obtaining a filtering image block corresponding to a target image block, the computer device can generate a second related image block for decoding a predicted image block in the multimedia data based on the filtering image block corresponding to the target image block, and store the second related image block in a decoding cache. Alternatively, the computer device can generate a display image block based on the filtering image block corresponding to the target image block, skip the step of storing the filtering image block corresponding to the target image block in a decoding cache, and transmit the display image block to a display device, which is configured to display the display image block.
[0084] In concrete implementation, as shown in Figure 13g, the neural network-based image filtering processor includes at least two convolutional layers, N Resblock layers, multiple activation layers, and one shuffle layer, where the at least two convolutional layers are denoted as convolutional layer 0, convolutional layer 1, convolutional layer 2, convolutional layer 3, ..., convolutional layer N, and the activation layers include activation layer 0, activation layer 1, activation layer 2, .... If the target coding / decoding information includes coding / decoding information corresponding to the first related image block and coding / decoding information corresponding to the target image block, the first related image block includes related image block 0 and related image block 1, the coding / decoding information corresponding to related image block 0 is coding / decoding information 0, and the coding / decoding information corresponding to related image block 1 is coding / decoding information 1. Encoded / decoded information 1 includes the image distance 1 between the associated image 1 and the target image block, and the reference direction 1 of the associated image 1 to the target image block. Encoded / decoded information 0 includes the image distance 0 between the associated image 0 and the target image block, and the reference direction 0 of the associated image 0 to the target image block. Convolutional layer 0 corresponds to associated image block 0, and convolutional layer 1 corresponds to associated image block 1. The encoded / decoded information corresponding to the target image block includes sequence-level quantization parameters, slice-level quantization parameters, slice-level coding type, block coding type, and a predicted image block corresponding to the target image block. The computer equipment can input the encoded / decoded information corresponding to the target image block as auxiliary information, along with the encoded / decoded information of the associated image block, into a neural network-based image filtering processor. Figure 13g shows one example, in which matrices are used to represent related image block 0, reference direction 0, and image distance 0, respectively. The convolutional layer 0 performs a convolution operation on the matrix corresponding to related image block 0, the matrix corresponding to reference direction 0, and the matrix corresponding to image distance 0 to obtain fused data corresponding to related image block 0.Matrices are used to represent related image block 1, reference direction 1, and image distance 1, respectively. A convolutional layer 1 performs a convolution operation on the matrix corresponding to related image block 1, the matrix corresponding to reference direction 1, and the matrix corresponding to image distance 1 to obtain fused data corresponding to related image block 1. The fused data corresponding to related image block 1 and the fused data corresponding to related image block 0 are determined as target fused data corresponding to the related image blocks. Simultaneously, matrices are used to represent sequence-level quantization parameters, slice-level quantization parameters, slice-level coding type, block coding type, and predicted image blocks corresponding to the target image blocks, respectively. These are input to a neural network-based image filtering processor, which performs a filtering process on the target image blocks based on the target fused data and target coding / decoding information to obtain filtered image blocks corresponding to the target image blocks.
[0085] It can be understood that, if the image to which the target image block belongs is an I-frame, a matrix with all elements being 0 is determined as the matrix corresponding to the slice coding type corresponding to the target image block, and if the image to which the target image block belongs is a B-frame or a P-frame, a matrix with all elements being 1 is determined as the matrix corresponding to the slice coding type corresponding to the target image block. The matrix corresponding to the block coding type of the target image block can be filled element by element according to the coding scheme of the coding unit. For example, if the coding unit belongs to an intra coding scheme, the value of the element corresponding to that coding unit is set to 0, and otherwise the value is set to 1. If the image to which the target image block belongs is an I-frame, related image block 0 and related image 1 can be filled using the target image block, and by setting image distance 1 and image distance 0 to the minimum weight of 0, a matrix corresponding to image distance 1 and a matrix corresponding to image distance 0 can be obtained. If the image to which the target image block belongs is a B-frame or a P-frame, set image distance 1 to (1025 - absolute value of the POC difference between related image block 1 and the target image block) / 1024 to obtain the matrix corresponding to image distance 1, and set image distance 0 to (1025 - absolute value of the POC difference between related image block 0 and the target image block) / 1024 to obtain the matrix corresponding to image distance 0. If the image to which the target image block belongs is an I-frame, set both the direction values (i.e., the elements of the matrix) corresponding to reference direction 0 and reference direction 1 to 0. If the image to which the target image block belongs is a B-frame or a P-frame, and the POC of the image to which related image block 0 belongs is smaller than the POC of the image to which the target image block belongs, set the direction value of reference direction 0 to 0; otherwise, set the direction value of reference direction 0 to 1. If the POC of the image to which related image block 1 belongs is smaller than the POC of the image to which the target image block belongs, set the direction value of reference direction 1 to 0; otherwise, set the direction value of reference direction 1 to 1.If the image to which the target image block belongs is a B-frame or a P-frame, related image block 0 may be obtained from the first reference frame in the reference frame list in the L0 direction, and related image block 1 may be obtained from the first reference frame in the reference frame list in the L1 direction. Both related image block 0 and related image block 1 have the same spatial positional relationship as the target image block, and the size of related image block 0 and related image block 1 are the same as the size of the target image block. The above processing method makes it possible for I-frames, B-frames, and P-frames to share the above-described neural network-based image filtering processor to perform filtering, saving resources and improving the versatility of the neural network-based image filtering processor by eliminating the need to train a separate neural network-based image filtering processor for each of the I-frames, B-frames, and P-frames.
[0086] In the embodiments of this application, a first related image block related to a target image block to be filtered within multimedia data is determined, target coding / decoding information related to the first related image block and the target image block is obtained, and a filtering process is performed on the target image block based on the target coding / decoding information and the first related image block to obtain a filtered image block corresponding to the target image block. By introducing multidimensional information such as related image blocks and target coding / decoding information, a richer amount of information is provided in the filtering process of the target image block, and at the same time, the target coding / decoding information is used to represent the degree of influence of the first related image block on the target image block. Therefore, the target coding / decoding information and the neural network-based image filtering processor make full use of the high-quality advantage of the related image block to perform filtering on the reconstructed image block, and by using related image blocks selectively in the filtering process for the target image block, it is possible to improve the accuracy of the filtering process of the target image block and improve the quality of the filtering process of the target image block. In other words, it is possible to improve the accuracy of the filtering process of images and improve the filtering effect of images, and furthermore, to improve the coding quality of multimedia data.
[0087] Referring to Figure 14, this is an exemplary structural diagram of a multimedia data processing device according to an embodiment of the present application. The multimedia data processing device may be a single computer program (including program code) executed on a computer device, for example, the multimedia data processing device may be a single application software, and the device may be configured to perform the corresponding steps in the method provided by the embodiment of the present application. As shown in Figure 14, the multimedia data processing device may comprise an acquisition module 141, a filtering module 142, a determination module 143, and a generation module 144.
[0088] The decision module is configured to determine a first related image block associated with a target image block to be filtered within multimedia data; the acquisition module is configured to acquire target coding / decoding information associated with the target image block; and the filtering module is configured to perform filtering on the target image block using a neural network-based image filtering processor based on the first related image block and the target coding / decoding information, thereby obtaining a filtered image block corresponding to the target image block.
[0089] As can be understood, the filtering module includes a fusion unit 14a and a filtering unit 15a, the fusion unit is configured to fuse the first related image block with the encoded / decoded information corresponding to the first related image block using the information fusion layer of a neural network-based image filtering processor to obtain target fused data corresponding to the first related image block, and the filtering unit is configured to perform filtering on the target image block based on the target fused data corresponding to the first related image block using the filtering processing layer of the neural network-based image filtering processor to obtain a filtered image block corresponding to the target image block.
[0090] In essence, the fusion unit is configured to fuse the first related image block with the encoded / decoded information corresponding to the first related image block using the information fusion layer of a neural network-based image filtering processor to obtain target fused data corresponding to the first related image block, and the filtering unit is configured to perform filtering on the target image block based on the target fused data corresponding to the first related image block and the encoded / decoded information corresponding to the target image block using the filtering processing layer of the neural network-based image filtering processor to obtain a filtered image block corresponding to the target image block.
[0091] It can be understood that the number of first related image blocks is M (where M is an integer greater than or equal to 1), the neural network-based image filtering processor includes M information fusion layers, each information fusion layer corresponds to one related image block, and the fusion unit, by the information fusion layer of the neural network-based image filtering processor, fuses the first related image block with the encoded / decoded information corresponding to the first related image block to obtain target fused data corresponding to the first related image block, which involves the information fusion layer of the neural network-based image filtering processor performing a convolution operation on the first related image block i (where i is a positive integer less than or equal to M) and the encoded / decoded information corresponding to the first related image block i to obtain fused data corresponding to the first related image block i, wherein the first related image block i belongs to one of the M first related image blocks, and when fused data corresponding to each of the M first related image blocks is obtained, the fused data corresponding to each of the M first related image blocks is determined as the target fused data corresponding to the first related image block.
[0092] It can be understood that, here, the number of first related image blocks is M (where M is an integer greater than or equal to 1), and the fusion unit fuses the first related image blocks with the encoded / decoded information corresponding to the first related image blocks using the information fusion layer of the neural network-based image filtering processor to obtain target fused data corresponding to the first related image blocks, which includes performing a convolution operation on the M first related image blocks and the encoded / decoded information using the information fusion layer of the neural network-based image filtering processor to obtain target fused data corresponding to the first related image blocks.
[0093] It can be understood that the number of first related image blocks is M (where M is an integer greater than or equal to 1), and the fusion unit fuses the first related image blocks with the encoded / decoded information corresponding to the first related image blocks using the information fusion layer of the neural network-based image filtering processor to obtain target fused data corresponding to the first related image blocks, which includes performing a dot product operation on the first related image block i (where i is a positive integer less than or equal to M) and the encoded / decoded information corresponding to the first related image block i using the information fusion layer of the neural network-based image filtering processor to obtain fused data corresponding to the first related image block i, wherein the first related image block i belongs to the M first related image blocks, and when fused data corresponding to each of the M first related image blocks is obtained, a convolution operation is performed on the fused data corresponding to each of the M first related image blocks to obtain target fused data corresponding to the first related image block.
[0094] It can be understood that the encoding / decoding information corresponding to the target image block includes at least one of sequence-level quantization parameters, slice-level quantization parameters, slice-level encoding type, block encoding type, filtering-enhanced image corresponding to the target image block, predicted image block information corresponding to the target image block, and block-divided image block information of the target image block. It can also be understood that the image block information corresponding to the target image block includes at least one of first color component information, second color component information, and third color component information, where the first color component information is used to represent the brightness of the image block, and the second and third color component information are used to represent the chroma of the image block. The image block information corresponding to the target image block includes at least one of filtering intensity image block information, block-divided image block information, and predicted image block information corresponding to the target image block.
[0095] As can be understood, the target coding / decoding information is used to represent the degree of influence of the first related image block on the target image block, and the target coding / decoding information includes at least one of the following: an influence factor of the first related image block on the target image block, a reference direction of the first related image block on the target image block, block division image block information of the first related image block, quantization parameters corresponding to the first related image block, reconstructed image block information corresponding to the first related image block, filtering intensity image block information corresponding to the first related image block, and predicted image block information corresponding to the first related image block, where the influence factor is determined based on the image distance between the first related image block and the target image.
[0096] The determination module is configured to determine the image distance between the first related image block and the target image block based on the image sequence count corresponding to the first related image block and the image sequence count corresponding to the target image block if the slice in which the target image block is located is an incomplete intra-encoded slice, and to determine a predetermined value as the image distance between the first related image block and the target image block if the slice in which the target image block is located is a complete intra-encoded slice.
[0097] As can be understood, the determination of the image distance between the first related image block and the target image block by the decision module, based on the image sequence count corresponding to the first related image block and the image sequence count corresponding to the target image block, includes determining the count difference between the image sequence count corresponding to the first related image block and the image sequence count corresponding to the target image block, and determining the absolute value of the count difference as the image distance between the first related image block and the target image block.
[0098] As can be understood, the determination of the image distance between the first related image block and the target image block by the decision module, based on the image sequence count corresponding to the first related image block and the image sequence count corresponding to the target image block, includes determining the count difference between the image sequence count corresponding to the first related image block and the image sequence count corresponding to the target image block, calculating a count mapping value corresponding to the absolute value of the count difference, and determining the image distance between the first related image block and the target image based on the count mapping value.
[0099] It can be understood that the image distance between the first related image block and the target image block is determined based on the time-domain layer of the target image block, or the image distance between the first related image block and the target image block is determined based on the level mapping value corresponding to the time-domain layer of the target image block.
[0100] It can be understood that the image block information corresponding to the first related image block includes at least one of the first color component information, the second color component information, and the third color component information, wherein the first color component information is used to represent the brightness of the image block, and the second and third color component information are used to represent the chroma of the image block, and the image block information corresponding to the first related image block includes one of the reconstructed image block information, block division image block information, filtering intensity image block information, and predicted image block information corresponding to the first related image block.
[0101] As can be understood, the decision module is further configured to determine the reference direction of the first associated image block to the target image block based on the relationship between the magnitude of the image sequence count corresponding to the first associated image block and the image sequence count corresponding to the target image block, if the slice in which the target image block is located is an incomplete intra-encoded slice, and to indicate the reference direction of the first associated image block to the target image block by adopting a predetermined direction value, if the slice in which the target image block is located is a complete intra-encoded slice.
[0102] As can be understood, the filtering module performs filtering on the target image block using a neural network-based image filtering processor based on the first related image block and the target coding / decoding information, and obtains a filtered image block corresponding to the target image block, which includes obtaining the image distance between the target image block and the first related image block from the target coding / decoding information; prohibiting filtering on the target image block based on the first related image block if the image distance is greater than a distance threshold; and performing filtering on the target image block using a neural network-based image filtering processor based on the first related image block and the target coding / decoding information if the image distance is less than or equal to the distance threshold, and obtaining a filtered image block corresponding to the target image block.
[0103] It can be understood that the filtering module, based on the first related image block and the target coding / decoding information, performs filtering on the target image block using a neural network-based image filtering processor to obtain a filtered image block corresponding to the target image block, which includes obtaining the image distance between the target image block and the first related image block from the target coding / decoding information; prohibiting filtering on the target image block based on the first related image block if the image distance does not fall within a specified image distance; and, if the image distance falls within a specified image distance, performing filtering on the target image block using a neural network-based image filtering processor based on the first related image block and the target coding / decoding information to obtain a filtered image block corresponding to the target image block.
[0104] It can be understood that the size of the target image block is the same as the size of the first related image block, or the size of the target image block is different from the size of the first related image block. The target image block and the first related image block have the same spatial positional relationship.
[0105] It can be understood that the size of the target image block is the same as the size of the first related image block, or the size of the target image block is different from the size of the first related image block, where there is a different spatial positional relationship between the target image block and the first related image block, and the size of the first related image block is determined based on the motion vector of the target image block.
[0106] It can be understood that, when the size of the target image block is smaller than the size of the first associated image block, the filtering module performs a filtering process on the target image block based on the first associated image block and the target coding / decoding information to obtain a filtered image block corresponding to the target image block, which includes dividing the first associated image block to obtain N image subblocks (where N is a positive integer greater than 1), wherein the target coding / decoding information includes coding / decoding information representing the degree of influence of each of the N image subblocks on the target image block, inputting the N image subblocks and the coding / decoding information corresponding to each of the N image subblocks into a neural network-based image filtering processor, and the neural network-based image filtering processor performs a filtering process on the target image block based on the N image subblocks and the coding / decoding information corresponding to each of the N image subblocks to obtain a filtered image block corresponding to the target image block.
[0107] As can be understood, the generation module is configured to generate a second related image block for decoding a predicted image block in the multimedia data based on a filtering image block corresponding to the target image block, and to store the second related image block in a decoding cache, or to generate a display image block based on a filtering image block corresponding to the root target image block, and to transmit the display image block to a display device, skipping the step of storing the filtering image block corresponding to the target image block in a decoding cache, and the display device is configured to display the display image block.
[0108] In the embodiments of this application, a first related image block related to a target image block to be filtered within multimedia data is determined, the first related image block and target coding / decoding information for the target image block are obtained, and a filtering process is performed on the target image block based on the target coding / decoding information and the first related image block to obtain a filtered image block corresponding to the target image block. By introducing multidimensional information such as related image blocks and target coding / decoding information, a richer amount of information is provided in the filtering process of the target image block. At the same time, the target coding / decoding information is used to represent the degree of influence of the first related image block on the target image block. Therefore, by using the target coding / decoding information to fully utilize the advantages of the high quality of the related image block when performing filtering on the reconstructed image block, the related image block is used selectively in the filtering process for the target image block, which helps to improve the accuracy of the filtering process of the target image block and improve the quality of the filtering process of the target image block. In other words, the accuracy of the image filtering process can be improved, the filtering effect of the image can be improved, and the coding quality of the multimedia data can be improved.
[0109] Referring to Figure 15, this is an exemplary structural diagram of a computer device according to an embodiment of the present application. As shown in Figure 15, the computer device 1000 may include a processor 1001, a network interface 1004, and memory 1005. In addition, the computer device 1000 may further include a user interface 1003 and at least one communication bus 1002. Here, the communication bus 1002 is configured to enable connection communication between these components. Here, the user interface 1003 may include a display and a keyboard, and optionally the user interface 1003 may further include a standard wired interface and a wireless interface. The network interface 1004 may optionally include a standard wired interface and a wireless interface (e.g., a Wi-Fi interface). The memory 1005 may be high-speed RAM memory or non-volatile memory, for example, at least one magnetic disk memory. Optionally, the memory 1005 may further include at least one storage device located away from the processor 1001 described above. As shown in Figure 15, the memory 1005, as a computer-readable storage medium, may include an operating system, a network communication module, a network interface module, and a device control application program.
[0110] In the computer device 1000 shown in Figure 15, the network interface 1004 can provide network communication functionality, and the user interface 1003 is configured to primarily provide an interface for inputting media content.
[0111] What can be understood is that the processor 1001 calls a device-specific application program stored in memory 1005, The steps include determining a first related image block related to the target image block to be filtered within the multimedia data, The steps include obtaining target coding and decoding information related to the first related image block and the target image block, The system can be configured to perform the following steps: based on the first related image block and the target coding / decoding information, a neural network-based image filtering processor performs a filtering process on the target image block to obtain a filtered image block corresponding to the target image block.
[0112] It can be understood that the processor 1001 may further be configured to call a device-specific application program stored in memory 1005 to perform a filtering process on the target image block using a neural network-based image filtering processor based on the first related image block and the target coding / decoding information, thereby obtaining a filtered image block corresponding to the target image block, and the above step may be: The steps include: using an information fusion layer of a neural network-based image filtering processor to fuse the first related image block with the encoded / decoded information corresponding to the first related image block to obtain target fused data corresponding to the first related image block; The process includes the step of using the filtering processing layer of the neural network-based image filtering processor to perform a filtering process on the target image block based on the target fusion data corresponding to the first related image block, thereby obtaining a filtered image block corresponding to the target image block.
[0113] It can be understood that the processor 1001 is configured to call a device-specific application program stored in memory 1005, and to perform filtering on the target image block using a neural network-based image filtering processor based on the first related image block and the target coding / decoding information, thereby obtaining a filtered image block corresponding to the target image block, and the above step is: The steps include: using an information fusion layer of a neural network-based image filtering processor to fuse the first related image block with the encoded / decoded information corresponding to the first related image block to obtain target fused data corresponding to the first related image block; The process includes the step of using the filtering processing layer of the neural network-based image filtering processor to perform a filtering process on the target image block based on the target fusion data corresponding to the first related image block and the encoding / decoding information corresponding to the target image block, thereby obtaining a filtered image block corresponding to the target image block.
[0114] It can be understood that the number of the first associated image blocks is M (where M is an integer greater than or equal to 1), and the neural network-based image filtering processor includes M information fusion layers, with each information fusion layer corresponding to one associated image block. The processor 1001 can be configured to call a device-specific application program stored in memory 1005 to perform the step of fusing the first related image block with the encoded / decoded information corresponding to the first related image block using the information fusion layer of a neural network-based image filtering processor to obtain target fused data corresponding to the first related image block, and the step is as follows: The steps include: performing a convolution operation on a first related image block i (where i is a positive integer less than or equal to M) and the encoded / decoded information corresponding to the first related image block i using the information fusion layer i of the neural network-based image filtering processor to obtain fused data corresponding to the first related image block i, wherein the first related image block i belongs to M first related image blocks, and the information fusion layer i is the information fusion layer corresponding to the first related image block i among the M information fusion layers; The method includes the step of determining, when fusion data corresponding to each of the M first related image blocks has been obtained, as the target fusion data corresponding to each of the M first related image blocks.
[0115] What can be understood is that the number of the aforementioned first related image blocks is M (where M is an integer greater than or equal to 1), The processor 1001 can be configured to call a device-specific application program stored in memory 1005 to perform the step of fusing the first related image block with the encoded / decoded information corresponding to the first related image block using the information fusion layer of a neural network-based image filtering processor to obtain target fused data corresponding to the first related image block, and the step is as follows: The process includes the step of performing a convolution operation on M first related image blocks and encoded / decoded information using the information fusion layer of the neural network-based image filtering processor to obtain target fused data corresponding to the first related image blocks.
[0116] It can be understood that the number of the first associated image blocks is M (where M is an integer greater than or equal to 1), and the neural network-based image filtering processor includes M information fusion layers, with each information fusion layer corresponding to one associated image block. The processor 1001 can be configured to call a device-specific application program stored in memory 1005 to perform the step of fusing the first related image block with the encoded / decoded information corresponding to the first related image block using the information fusion layer of a neural network-based image filtering processor to obtain target fused data corresponding to the first related image block, and the step is as follows: The step of performing a dot product operation on a first related image block i (where i is a positive integer less than or equal to M) and the encoded / decoded information corresponding to the first related image block i using the information fusion layer i of the neural network-based image filtering processor, thereby obtaining fused data corresponding to the first related image block i, wherein the first related image block i belongs to M first related image blocks, and the information fusion layer i is the information fusion layer corresponding to the first related image block i among the M information fusion layers, The method includes the step of obtaining fused data corresponding to each of the M first related image blocks, performing a convolution operation on the fused data corresponding to each of the M first related image blocks, and obtaining target fused data corresponding to the first related image block.
[0117] It can be understood that the coding and decoding information corresponding to the target image block includes at least one of the following: sequence-level quantization parameters, slice-level quantization parameters, slice-level coding type, block coding type, filtering intensity image block information corresponding to the target image block, predicted image block information corresponding to the target image block, and block division image block information of the target image block.
[0118] It can be understood that the image block information corresponding to the target image block includes at least one of the first color component information, the second color component information, and the third color component information, where the first color component information is used to represent the brightness of the image block, and the second and third color component information are used to represent the chroma of the image block, and the image block information corresponding to the target image block includes one of the filtering intensity image block information, block division image block information, and predicted image block information corresponding to the target image block.
[0119] It can be understood that the target coding / decoding information is used to represent the degree of influence of the first related image block on the target image block, and the target coding / decoding information includes at least one of the following: an influence factor of the first related image block on the target image block, a reference direction of the first related image block on the target image block, block division image block information of the first related image block, quantization parameters corresponding to the first related image block, reconstructed image block information corresponding to the first related image block, filtering intensity image block information corresponding to the first related image block, and predicted image block information corresponding to the first related image block, where the influence factor is determined based on the image distance between the first related image block and the target image block.
[0120] What can be understood is that the processor 1001 further calls a device-specific application program stored in memory 1005 to determine the image distance between the first related image block and the target image block, based on the image sequence count corresponding to the first related image block and the image sequence count corresponding to the target image block, if the slice in which the target image block is located is an incomplete intra-encoded slice. The system can be configured to perform the following steps: if the slice in which the target image block is located is a fully intra-encoded slice, determine a predetermined distance value as the image distance between the first associated image block and the target image block.
[0121] It can be understood that the processor 1001 may be configured to call a device-specific application program stored in memory 1005 to determine the image distance between the first associated image block and the target image block based on the image sequence count corresponding to the first associated image block and the image sequence count corresponding to the target image block, and the step is as follows: The steps include determining the difference between the image sequence count corresponding to the first related image block and the image sequence count corresponding to the target image block, The process includes the step of determining the absolute value of the count difference as the image distance between the first related image block and the target image block.
[0122] It can be understood that the processor 1001 may be configured to call a device-specific application program stored in memory 1005 to determine the image distance between the first associated image block and the target image block based on the image sequence count corresponding to the first associated image block and the image sequence count corresponding to the target image block, and the step is as follows: The steps include determining the difference between the image sequence count corresponding to the first related image block and the image sequence count corresponding to the target image block, The process includes the steps of calculating a count mapping value corresponding to the absolute value of the count difference, and determining the image distance between the first related image block and the target image block based on the count mapping value.
[0123] It can be understood that the image distance between the first related image block and the target image block is determined based on the time-domain layer of the target image block, or the image distance between the first related image block and the target image block is determined based on the level mapping value corresponding to the time-domain layer of the target image block.
[0124] It can be understood that the image block information corresponding to the first related image block includes at least one of the first color component information, the second color component information, and the third color component information, wherein the first color component information is used to represent the brightness of the image block, and the second and third color component information are used to represent the chroma of the image block, and the image block information corresponding to the first related image block includes one of the reconstructed image block information, block division image block information, filtering intensity image block information, and predicted image block information corresponding to the first related image block.
[0125] What can be understood is that the processor 1001 calls a device-specific application program stored in memory 1005 to determine the reference direction of the first associated image block to the target image block, based on the relationship between the magnitude of the image sequence count corresponding to the first associated image block and the image sequence count corresponding to the target image block, if the slice in which the target image block is located is an incomplete intra-encoded slice. If the slice in which the target image block is located is a fully intra-encoded slice, the system can be configured to perform the steps of: adopting a predetermined direction value to indicate the reference direction of the first associated image block to the target image block.
[0126] It can be understood that the processor 1001 is configured to call a device-specific application program stored in memory 1005, and to perform filtering on the target image block using a neural network-based image filtering processor based on the first related image block and the target coding / decoding information, thereby obtaining a filtered image block corresponding to the target image block, and the above step is: The steps include obtaining the image distance between the target image block and the first related image block from the target coding and decoding information, If the image distance is greater than the distance threshold, the step of prohibiting filtering the target image block based on the first associated image block, If the image distance is less than or equal to a distance threshold, the process includes the step of performing a filtering operation on the target image block using a neural network-based image filtering processor based on the first related image block and the target coding / decoding information to obtain a filtered image block corresponding to the target image block.
[0127] It can be understood that the processor 1001 is configured to call a device-specific application program stored in memory 1005, and to perform filtering on the target image block using a neural network-based image filtering processor based on the first related image block and the target coding / decoding information, thereby obtaining a filtered image block corresponding to the target image block, and the above step is: The steps include obtaining the image distance between the target image block and the first related image block from the target coding and decoding information, If the aforementioned image distance does not fall within the specified image distance, the step of prohibiting filtering the target image block based on the first associated image block, If the image distance falls within a specified image distance, the process includes the step of performing a filtering operation on the target image block using a neural network-based image filtering processor based on the first related image block and the target coding / decoding information, thereby obtaining a filtered image block corresponding to the target image block.
[0128] It can be understood that the size of the target image block is the same as the size of the first related image block, or the size of the target image block is different from the size of the first related image block. The target image block and the first related image block have the same spatial positional relationship.
[0129] It can be understood that the size of the target image block is the same as the size of the first related image block, or the size of the target image block is different from the size of the first related image block. There is a different spatial positional relationship between the target image block and the first related image block, and the first related image block is determined based on the motion vector of the target image block.
[0130] It can be understood that the processor 1001 is configured to call a device-specific application program stored in memory 1005 to perform the following steps: if the size of the target image block is smaller than the size of the first related image block, the neural network-based image filtering processor performs filtering on the target image block based on the first related image block and the target coding / decoding information, thereby obtaining a filtered image block corresponding to the target image block; and the above steps are performed as follows: A step of dividing the first associated image block to obtain N image subblocks (where N is a positive integer greater than 1), wherein the target coding / decoding information includes coding / decoding information for representing the degree of influence of each of the N image subblocks on the target image block. The steps include inputting the N image subblocks and the encoding / decoding information corresponding to each of the N image subblocks into an image filtering processor based on a neural network, The process includes the step of using the neural network-based image filtering processor to perform a filtering process on the target image block based on the N image subblocks and the encoding / decoding information corresponding to each of the N image subblocks, thereby obtaining a filtered image block corresponding to the target image block.
[0131] What can be understood is that the processor 1001 calls a device-specific application program stored in memory 1005, A step of generating a second related image block for decoding a predicted image block in the multimedia data based on a filtering image block corresponding to the target image block, and storing the second related image block in a decoding cache, or The system can be configured to generate a display image block based on a filtering image block corresponding to the target image block, skip the step of storing the filtering image block corresponding to the target image block in a decoding cache, and transmit the display image block to a display device, the display device being configured to display the display image block.
[0132] In the embodiments of this application, a first related image block related to a target image block to be filtered within multimedia data is determined, target coding / decoding information of the first related image block and the degree of influence of the target image block is obtained, and a filtering process is performed on the target image block based on the target coding / decoding information and the first related image block to obtain a filtered image block corresponding to the target image block. By introducing multidimensional information such as related image blocks and target coding / decoding information, a richer amount of information is provided in the filtering process of the target image block, and at the same time, the target coding / decoding information is used to represent the degree of influence of the first related image block on the target image block. Therefore, the target coding / decoding information and a neural network-based image filtering processor make full use of the advantages of the high quality of the related image block to perform filtering on the reconstructed image block, and by using related image blocks selectively in the filtering process for the target image block, it is possible to improve the accuracy of the filtering process of the target image block and improve the quality of the filtering process of the target image block, that is, to improve the accuracy of the filtering process of the image and improve the filtering effect of the image, and furthermore, to improve the coding quality of the multimedia data.
[0133] It should be understood that the computer device 1000 described in the embodiments of this application can perform the above-described multimedia data processing method in the embodiment corresponding to Figure 4, and can also perform the above-described multimedia data processing device in the embodiment corresponding to Figure 14, which will not be repeated here. Furthermore, the beneficial effects of employing the same method will not be repeated here.
[0134] Furthermore, it should be noted that the embodiments of this application further provide a computer-readable storage medium, the computer-readable storage medium storing a computer program executed by the multimedia data processing device referred to in the above specification, the computer program including program instructions, and when the processor executes the program instructions, the above description of the multimedia data processing method in the embodiment corresponding to Figure 4 can be performed, and will not be repeated here. Furthermore, the beneficial effects of employing the same method will not be repeated. For technical details not disclosed in the embodiments of the computer-readable storage medium of this application, please refer to the description of the method embodiments of this application.
[0135] For example, the above program instructions may be deployed to run on one computer device, or on at least two computer devices located in the same place, or on at least two computer devices distributed to at least two locations and interconnected by a communication network, and at least two computer devices distributed to at least two locations and interconnected by a communication network can constitute a blockchain network.
[0136] The computer-readable storage medium described above may be the hardware or memory of the data processing device or computer equipment according to any of the embodiments described above, or it may be an internal storage unit of the computer equipment. The computer-readable storage medium may also be an external storage device of the computer equipment, such as a plug-in hard disk, SmartMedia card (SMC), Secure Digital (SD) card, or flash card, which is installed in the computer equipment. Furthermore, the computer-readable storage medium may further include both the internal storage unit and the external storage device of the computer equipment. The computer-readable storage medium is configured to store the computer program and other programs and data necessary for the computer equipment. The computer-readable storage medium may further be configured to temporarily store output or output data.
[0137] The embodiments of this application further provide a computer program product including a computer program / instruction, which, when executed by a processor, causes the processor to implement the above-described multimedia data processing method in the embodiment corresponding to Figure 4, and therefore will not be repeated here. Furthermore, the beneficial effects of employing the same method will not be repeated. For technical details not disclosed in the embodiments of the computer program product of this application, please refer to the description of the method embodiment of this application.
[0138] The terms “First,” “Second,” and “Third,” etc., in the specification, claims, and accompanying drawings of the embodiments of this application are not intended to limit a specific order but to distinguish different media content. Furthermore, “Includes” and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, apparatus, product, or device comprising a series of steps or units is not limited to the listed steps or modules, but may also include, exemplarily, steps or modules not listed, or other step units specific to these processes, methods, apparatus, products, or devices.
[0139] As will be obvious to those skilled in the art, the units and algorithmic steps of each embodiment described with reference to the embodiments disclosed herein are implemented by electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the above description uses general terms to describe the configuration and steps of each embodiment according to their function. Whether these functions are performed in hardware form or software form depends on the specific application and design constraints of the technical solution. A person skilled in the art may implement the described functions using different methods depending on each specific application, but such implementations should not be considered to exceed the scope of protection of this application.
[0140] The methods and related apparatus provided by embodiments of this application are described with reference to flowcharts and / or exemplary structural diagrams of the methods according to embodiments of this application, specifically, computer program instructions may be used to realize each process and / or block in the flowchart and / or exemplary structural diagram of the method, and combinations of processes and / or blocks in the flowchart and / or block diagram. To generate a machine, these computer program instructions are provided to the processor of a general-purpose computer, a dedicated computer, an embedded processor, or another programmable data processing device, causing the instructions executed by the processor of the computer or other programmable data processing device to generate an apparatus for performing the functions specified in one or more processes in the flowchart and / or one or more blocks in the exemplary structural diagram. These computer program instructions may also be stored in computer-readable memory that operates the computer or other programmable data processing device in a particular manner, and the instructions stored in said computer-readable memory generate a product including instruction means for realizing the functions specified in one or more processes in the flowchart and / or one or more blocks in the exemplary structural diagram. These computer program instructions can also be loaded into a computer or other programmable data processing device to cause the computer or other programmable device to perform a series of operational steps to generate processing to be implemented by the computer, thereby providing steps for realizing one or more processes in a flowchart and / or one or more blocks in an exemplary structural diagram.
[0141] The above disclosures represent only preferred embodiments of the present application and do not limit the scope of the claims; therefore, the same modifications made in accordance with the claims of the present application are still within the scope of the present application.
Claims
1. A multimedia data processing method performed by a computer device, The steps include determining a first related image block related to the target image block to be filtered within the multimedia data, The steps include: obtaining target coding and decoding information related to the aforementioned target image block; A step of performing a filtering process on the target image block using a neural network-based image filtering processor based on the first related image block and the target coding / decoding information, in order to obtain a filtered image block corresponding to the target image block, The steps include: using the information fusion layer of the neural network-based image filtering processor to fuse the first related image block with the encoded / decoded information corresponding to the first related image block to obtain target fused data corresponding to the first related image block; The process includes the step of performing a filtering process on the target image block based on the target fusion data corresponding to the first related image block using the filtering processing layer of the neural network-based image filtering processor, thereby obtaining a filtered image block corresponding to the target image block. A multimedia data processing method including [specific data processing method].
2. A multimedia data processing method performed by a computer device, The steps include determining a first related image block related to the target image block to be filtered within the multimedia data, The steps include: obtaining target coding and decoding information related to the aforementioned target image block; A step of performing a filtering process on the target image block using a neural network-based image filtering processor based on the first related image block and the target coding / decoding information, in order to obtain a filtered image block corresponding to the target image block, The steps include: using the information fusion layer of the neural network-based image filtering processor to fuse the first related image block with the encoded / decoded information corresponding to the first related image block to obtain target fused data corresponding to the first related image block; The process includes the step of using the filtering processing layer of the neural network-based image filtering processor to perform a filtering process on the target image block based on the target fusion data corresponding to the first related image block and the encoded / decoded information corresponding to the target image block, thereby obtaining a filtered image block corresponding to the target image block. A multimedia data processing method including [specific data processing method].
3. The number of the first related image blocks is M, where M is an integer greater than or equal to 1, and the neural network-based image filtering processor includes M information fusion layers, where one information fusion layer corresponds to one related image block. The step of obtaining target fused data corresponding to the first related image block by fusing the first related image block with the encoded / decoded information corresponding to the first related image block using the information fusion layer of the neural network-based image filtering processor is as follows: The steps include: performing a convolution operation on a first related image block i and encoded / decoded information corresponding to the first related image block i using an information fusion layer i of the neural network-based image filtering processor to obtain fused data corresponding to the first related image block i, wherein the first related image block i belongs to M first related image blocks, i is a positive integer less than or equal to M, and the information fusion layer i is the information fusion layer corresponding to the first related image block i among the M information fusion layers; The process includes the step of obtaining fusion data corresponding to each of the M first related image blocks, and then determining the fusion data corresponding to each of the M first related image blocks as the target fusion data corresponding to the first related image block. The multimedia data processing method according to claim 1.
4. The number of the first related image blocks is M, where M is an integer greater than or equal to 1. The step of obtaining target fused data corresponding to the first related image block by fusing the first related image block with the encoded / decoded information corresponding to the first related image block using the information fusion layer of the neural network-based image filtering processor is as follows: The process includes the step of performing a convolution operation on M first related image blocks and encoded / decoded information using the information fusion layer of the neural network-based image filtering processor to obtain target fused data corresponding to the first related image blocks. The multimedia data processing method according to claim 1.
5. The number of the first related image blocks is M, where M is an integer greater than or equal to 1, and the neural network-based image filtering processor includes M information fusion layers, where one information fusion layer corresponds to one related image block. The step of obtaining target fused data corresponding to the first related image block by fusing the first related image block with the encoded / decoded information corresponding to the first related image block using the information fusion layer of a neural network-based image filtering processor is as follows: A step of performing a dot product operation on a first related image block i and encoded / decoded information corresponding to the first related image block i using an information fusion layer i of an image filtering processor based on a neural network, thereby obtaining fused data corresponding to the first related image block i, wherein the first related image block i belongs to M first related image blocks, i is a positive integer less than or equal to M, and the information fusion layer i is the information fusion layer corresponding to the first related image block i among the M information fusion layers. The process includes the step of obtaining fused data corresponding to each of the M first related image blocks, performing a convolution operation on the fused data corresponding to each of the M first related image blocks, and obtaining target fused data corresponding to the first related image block. The multimedia data processing method according to claim 1.
6. The encoding and decoding information corresponding to the target image block includes at least one of sequence-level quantization parameters, slice-level quantization parameters, slice-level encoding type, block encoding type, filtering intensity image block information corresponding to the target image block, prediction image block information corresponding to the target image block, and block division image block information of the target image block. The multimedia data processing method according to claim 2.
7. The image block information corresponding to the target image block includes at least one of the first color component information, the second color component information, and the third color component information. The first color component information is used to represent the brightness of the image block, the second color component information and the third color component information are used to represent the chroma of the image block, and the image block information corresponding to the target image block includes one of the filtering intensity image block information, block division image block information, and predicted image block information corresponding to the target image block. The multimedia data processing method according to claim 6.
8. The target encoding and decoding information is used to represent the degree of influence of the first related image block on the target image block. The target coding and decoding information includes at least one of the following: an influence factor of the first related image block on the target image block, a reference direction of the first related image block on the target image block, block division image block information of the first related image block, quantization parameters corresponding to the first related image block, reconstructed image block information corresponding to the first related image block, filtering intensity image block information corresponding to the first related image block, and predicted image block information corresponding to the first related image block. The influencing factor is determined based on the image distance between the first related image block and the target image block. The multimedia data processing method according to claim 1.
9. The aforementioned multimedia data processing method is If the slice in which the target image block is located is an incomplete intra-encoded slice, the step of determining the image distance between the first related image block and the target image block based on the image sequence count corresponding to the first related image block and the image sequence count corresponding to the target image block, If the slice in which the target image block is located is a fully intra-encoded slice, the further step is to determine a predetermined distance value as the image distance between the first associated image block and the target image block. The multimedia data processing method according to claim 8.
10. The step of determining the image distance between the first related image block and the target image block based on the image sequence count corresponding to the first related image block and the image sequence count corresponding to the target image block is: The steps include determining the difference between the image sequence count corresponding to the first related image block and the image sequence count corresponding to the target image block, The step includes determining the absolute value of the count difference as the image distance between the first related image block and the target image block, The multimedia data processing method according to claim 9.
11. The step of determining the image distance between the first related image block and the target image block based on the image sequence count corresponding to the first related image block and the image sequence count corresponding to the target image block is: The steps include determining the difference between the image sequence count corresponding to the first related image block and the image sequence count corresponding to the target image block, The steps include: calculating a count mapping value corresponding to the absolute value of the count difference, and determining the image distance between the first related image block and the target image block based on the count mapping value; The multimedia data processing method according to claim 9.
12. The image distance between the first related image block and the target image block is determined based on the time-domain layer of the target image block, or the image distance between the first related image block and the target image block is determined based on the level mapping value corresponding to the time-domain layer of the target image block. The multimedia data processing method according to claim 8.
13. The image block information corresponding to the first related image block includes at least one of the first color component information, the second color component information, and the third color component information. The first color component information is used to represent the brightness of the image block, the second color component information and the third color component information are used to represent the chroma of the image block, and the image block information corresponding to the first associated image block includes one of the reconstructed image block information, block division image block information, filtering intensity image block information, and predicted image block information corresponding to the first associated image block. The multimedia data processing method according to claim 8.
14. The aforementioned multimedia data processing method is If the slice in which the target image block is located is an incomplete intra-encoded slice, the steps include determining the reference direction of the first associated image block to the target image block based on the relationship between the magnitude of the image sequence count corresponding to the first associated image block and the image sequence count corresponding to the target image block, If the slice in which the target image block is located is a fully intra-encoded slice, the further step includes taking a predetermined direction value to indicate the reference direction of the first associated image block to the target image block. The multimedia data processing method according to claim 8.
15. A multimedia data processing method performed by a computer device, The steps include determining a first related image block related to the target image block to be filtered within the multimedia data, The steps include: obtaining target coding and decoding information related to the aforementioned target image block; A step of performing a filtering process on the target image block using a neural network-based image filtering processor based on the first related image block and the target coding / decoding information, in order to obtain a filtered image block corresponding to the target image block, The steps include obtaining the image distance between the target image block and the first related image block from the target coding and decoding information, If the image distance is greater than the distance threshold, the step of prohibiting filtering the target image block based on the first associated image block, The steps include: if the image distance is less than or equal to a distance threshold, a neural network-based image filtering processor performs a filtering process on the target image block based on the first related image block and the target coding / decoding information to obtain a filtered image block corresponding to the target image block; A multimedia data processing method including [specific data processing method].
16. A multimedia data processing method performed by a computer device, The steps include determining a first related image block related to the target image block to be filtered within the multimedia data, The steps include: obtaining target coding and decoding information related to the aforementioned target image block; A step of performing a filtering process on the target image block using a neural network-based image filtering processor based on the first related image block and the target coding / decoding information, in order to obtain a filtered image block corresponding to the target image block, The steps include obtaining the image distance between the target image block and the first related image block from the target coding and decoding information, If the aforementioned image distance does not fall within the specified image distance, the step of prohibiting filtering the target image block based on the first related image block, The steps include: if the aforementioned image distance falls within a specified image distance, a neural network-based image filtering processor performs a filtering process on the target image block based on the first related image block and the target coding / decoding information to obtain a filtered image block corresponding to the target image block; A multimedia data processing method including [specific data processing method].
17. A multimedia data processing method performed by a computer device, The steps include determining a first related image block related to the target image block to be filtered within the multimedia data, The steps include: obtaining target coding and decoding information related to the aforementioned target image block; If the size of the target image block is smaller than the size of the first related image block, the image filtering processor based on a neural network performs a filtering process on the target image block based on the first related image block and the target coding / decoding information to obtain a filtered image block corresponding to the target image block, A step of dividing the first related image block to obtain N image subblocks, wherein N is a positive integer greater than 1, and the target encoding / decoding information includes encoding / decoding information for representing the degree of influence of each of the N image subblocks on the target image block. The steps include inputting the N image subblocks and the encoding / decoding information corresponding to each of the N image subblocks into an image filtering processor based on a neural network, The process includes the step of using the neural network-based image filtering processor to perform a filtering process on the target image block based on the N image subblocks and the encoding / decoding information corresponding to each of the N image subblocks, thereby obtaining a filtered image block corresponding to the target image block. A multimedia data processing method comprising the following, wherein the target image block and the first related image block have the same spatial positional relationship.
18. A multimedia data processing method performed by a computer device, The steps include determining a first related image block related to the target image block to be filtered within the multimedia data, The steps include: obtaining target coding and decoding information related to the aforementioned target image block; The steps include: obtaining target coding and decoding information related to the aforementioned target image block; The process includes the step of performing a filtering process on the target image block using a neural network-based image filtering processor based on the first related image block and the target coding / decoding information, to obtain a filtered image block corresponding to the target image block. The aforementioned multimedia data processing method is A step of generating a second related image block for decoding a predicted image block in the multimedia data based on a filtering image block corresponding to the target image block, and storing the second related image block in a decoding cache, or A step of generating a display image block based on a filtering image block corresponding to the target image block, and transmitting the display image block to a display device, skipping the step of storing the filtering image block corresponding to the target image block in a decoding cache, wherein the display device is configured to display the display image block. A method for processing multimedia data.
19. A multimedia data processing device, A decision module configured to determine a first associated image block related to a target image block to be filtered within multimedia data, An acquisition module configured to acquire target coding and decoding information related to the first related image block and the target image block, The system includes a filtering module configured to perform filtering on the target image block using a neural network-based image filtering processor based on the first related image block and the target coding / decoding information, thereby obtaining a filtered image block corresponding to the target image block. The filtering module described above is The information fusion layer of the neural network-based image filtering processor fuses the first related image block with the encoded / decoded information corresponding to the first related image block to obtain target fused data corresponding to the first related image block. The filtering layer of the neural network-based image filtering processor performs filtering on the target image block based on the target fusion data corresponding to the first related image block, thereby obtaining a filtered image block corresponding to the target image block. A multimedia data processing device further configured as follows.
20. A computer device comprising a processor and memory, A computer device comprising a processor connected to memory, the memory configured to store program code, and the processor configured to execute the multimedia data processing method described in any one of claims 1 to 18 by calling the program code.
21. A computer program that causes a processor to execute the multimedia data processing method described in any one of claims 1 to 18.
Citation Information
Patent Citations
Encoding program, decoding program, encoding device, decoding device, encoding method, and decoding method
JP2020198463A
FILTERING METHOD, APPARATUS, ENCODER AND COMPUTER STORAGE MEDIUM
JP2022526107A
Multiple Neural Network Models for Filtering During Video Coding
JP2024501331A
Method and apparatus for filtering with mode-aware deep learning
US20200213587A1
Multiple neural network models for filtering during video coding
US20220215593A1