Image processing method, processing equipment and storage medium

The image blocks are filtered through multiple input branches and neural networks or lookup tables, which solves the calculation complexity problem in the loop filtering transformation stage of neural networks and improves the video encoding and decoding efficiency.

CN120455665AActive Publication Date: 2025-08-08SHENZHEN TRANSSION HLDG CO LTD

Patent Information

Application Number
CN202510812265.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-17
Publication Date
2025-08-08
Estimated Expiration
2045-06-17

AI Technical Summary

Technical Problem

In the existing high-efficiency video encoding framework, the transformation stage of neural network loop filtering increases the number of convolution operations and increases the computational complexity, which limits the video encoding and decoding efficiency.

Method used

Multi-input branches and neural networks or lookup tables are used to filter the image blocks, and the complexity of the filtering process is reduced through channel stitching and packet convolution modules.

Benefits of technology

The complexity of filtering processing is reduced and the efficiency of video encoding and decoding is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120455665A_ABST
    Figure CN120455665A_ABST
Patent Text Reader

Abstract

The invention provides an image processing method, processing equipment and a storage medium. The image processing method comprises the following steps: filtering at least one image block according to multiple input branches, a neural network and / or a lookup table. According to the technical scheme of the invention, the complexity of filtering processing can be reduced, and the video coding and / or decoding efficiency can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image processing technology, and in particular to an image processing method, processing device and storage medium. Background Art

[0002] Existing high-efficiency video coding frameworks, such as Neural Network Based Video Coding (NNVC) and / or Enhanced Compression Model (ECM), propose a video frame encoding technology to improve coding performance without significantly increasing computational complexity.

[0003] During the process of conceiving and implementing this application, the inventors discovered that there are at least the following problems: during the transformation stage of the neural network loop filtering during the encoding and decoding process, due to the diversity of image block features, the number of convolution operations of the neural network increases, the computational complexity increases, and thus limits the efficiency of video encoding and / or decoding.

[0004] The preceding description is intended to provide general background information and does not necessarily constitute prior art. Summary of the Invention

[0005] In response to the above technical problems, the present application provides an image processing method, a processing device and a storage medium, aiming to solve the technical problem of how to reduce the complexity of filtering processing and thereby support improving the efficiency of video encoding and / or decoding.

[0006] The present application provides an image processing method, which can be applied to a processing device, comprising the steps of:

[0007] S1, performing filtering processing on at least one image block according to multiple input branches, a neural network and / or a lookup table.

[0008] Optionally, the image block feature comprises at least one of a reconstruction transformation feature of a reconstructed block of at least one image block, a prediction transformation feature of a prediction block of at least one image block, and a derivative transformation feature of a derivative block of at least one image block;

[0009] Optionally, step S1 includes the steps of:

[0010] S11, performing channel splicing based on multiple input branches and at least one image block feature to determine or obtain input features;

[0011] S12, performing filtering processing on at least one input feature according to a neural network and / or a lookup table.

[0012] Optionally, the multi-input branch includes at least one of the following:

[0013] a first input branch for processing a luminance component;

[0014] a second input branch for processing chrominance components;

[0015] The third input branch is used for mixing the luminance component and the chrominance component.

[0016] Optionally, the first input branch includes at least one of the following: a first sub-input branch for processing the reconstructed transformation features and the derived transformation features on the luminance component, a second sub-input branch for processing the predicted transformation features and the derived transformation features on the luminance component, a third sub-input branch for processing the reconstructed transformation features on the luminance component, a fourth sub-input branch for processing the predicted transformation features on the luminance component, and a fifth sub-input branch for processing the derived transformation features on the luminance component.

[0017] Optionally, the second input branch includes at least one of the following: a sixth sub-input branch for processing the reconstructed transformation features and the derived transformation features on the chroma component, a seventh sub-input branch for processing the predicted transformation features and the derived transformation features on the chroma component, an eighth sub-input branch for processing the reconstructed transformation features on the chroma component, a ninth sub-input branch for processing the predicted transformation features on the chroma component, a tenth sub-input branch for processing the derived transformation features on the chroma component, a fourth input branch for processing the U component, and a fifth input branch for processing the V component.

[0018] Optionally, performing channel stitching based on the multiple input branches and at least one image block feature includes at least one of the following:

[0019] performing channel concatenation on the predicted transformation feature of at least one luminance component, the reconstructed transformation feature of at least one luminance component, and the derived transformation feature of at least one luminance component according to the first input branch;

[0020] performing channel concatenation on the predicted transform feature of at least one chroma component, the reconstructed transform feature of at least one chroma component, and the derived transform feature of at least one luminance component according to the second input branch;

[0021] performing channel splicing on at least one of a predicted transform feature of at least one luminance component and / or chrominance component, a derived transform feature of at least one luminance component and / or chrominance component, and a reconstructed transform feature of at least one luminance component and / or chrominance component according to the third input branch;

[0022] performing channel splicing on the reconstructed transformation feature of at least one luminance component and the derived transformation feature of at least one luminance component according to the first sub-input branch;

[0023] performing channel concatenation on the predicted transformation feature of at least one luminance component and the derived transformation feature of at least one luminance component according to the second sub-input branch;

[0024] performing channel splicing on the reconstructed transformation features of at least one brightness component according to the third sub-input branch;

[0025] performing channel splicing on the predicted transformation feature of at least one luminance component according to the fourth sub-input branch;

[0026] performing channel splicing on the derived transformation features of at least one luminance component according to the fifth sub-input branch;

[0027] performing channel splicing on the reconstructed transformation feature of at least one chroma component and the derived transformation feature of at least one chroma component according to the sixth sub-input branch;

[0028] performing channel concatenation on the predicted transformation feature of at least one chroma component and the derived transformation feature of at least one chroma component according to the seventh sub-input branch;

[0029] performing channel splicing on the reconstructed transformation feature of at least one chroma component according to the eighth sub-input branch;

[0030] performing channel splicing on the predicted transformation feature of at least one chrominance component according to the ninth sub-input branch;

[0031] performing channel splicing on the derived transform feature of at least one chroma component according to the tenth sub-input branch;

[0032] performing channel concatenation on the predicted transformation feature of at least one U component, the reconstructed transformation feature of at least one U component, and the derived transformation feature of at least one U component according to the fourth input branch;

[0033] According to the fifth input branch, channel concatenation is performed on the predicted transformation feature of the at least one V component, the reconstructed transformation feature of the at least one V component, and the derived transformation feature of the at least one V component.

[0034] Optionally, the neural network includes a grouped convolution module for dividing at least one input feature into at least one group for convolution processing.

[0035] Optionally, at least one group includes at least one of the following:

[0036] for processing at least one group of features corresponding to the first input branch in at least one input feature;

[0037] for processing at least one group of features corresponding to the second input branch in at least one input feature;

[0038] for processing at least one group of features corresponding to a third input branch in at least one input feature;

[0039] for processing at least one group of features corresponding to the first sub-input branch in at least one input feature;

[0040] for processing at least one group of features corresponding to the second sub-input branch in at least one input feature;

[0041] for processing at least one group of features corresponding to the third sub-input branch in at least one input feature;

[0042] for processing at least one group of features corresponding to a fourth sub-input branch in at least one input feature;

[0043] for processing at least one group of features corresponding to the fifth sub-input branch in at least one input feature;

[0044] for processing at least one group of features corresponding to a sixth sub-input branch in at least one input feature;

[0045] for processing at least one group of features corresponding to a seventh sub-input branch in at least one input feature;

[0046] for processing at least one group of features corresponding to an eighth sub-input branch in at least one input feature;

[0047] for processing at least one group of features corresponding to a ninth sub-input branch in at least one input feature;

[0048] for processing at least one group of features corresponding to a tenth sub-input branch in at least one input feature;

[0049] for processing at least one group of features corresponding to a fourth input branch in at least one input feature;

[0050] for processing at least one group of features corresponding to a fifth input branch in at least one input feature;

[0051] At least one group includes at least one channel of the plurality of channels of each input branch.

[0052] Optionally, the input branch is at least one item in multiple input branches corresponding to at least one input feature.

[0053] Optionally, the image processing method further includes at least one of the following:

[0054] The grouped convolution module includes at least one convolution module corresponding to a group, and the convolution modules corresponding to at least one group are independent of each other;

[0055] At least two groups contain the same number of input channels;

[0056] The convolutional layer parameters of at least two groups are the same;

[0057] The grouped convolution module sequentially includes the precursor layer of each group, the first channel shuffle layer, and the subsequent convolution module of each group;

[0058] The input of the first channel shuffle layer is the output of at least two groups of predecessor layers, and the output of the first channel shuffle layer is the input of at least two subsequent groups of convolution modules;

[0059] The precursor layer includes at least one convolution module for performing convolution processing on at least one set of features;

[0060] The first channel shuffling layer is used to permute the output feature maps of the predecessor layer across groups in the channel dimension through the first channel rearrangement rule;

[0061] The grouped convolution module includes a second channel shuffle layer and a convolution module for each group, and the input of the convolution module for each group is the output of the second channel shuffle layer;

[0062] The second channel shuffling layer is used to perform cross-group permutation on the channel dimension of the input feature map including features of at least two groups through a second channel rearrangement rule;

[0063] The grouped convolution module sequentially includes the second channel shuffle layer, the predecessor layer of each group, the first channel shuffle layer, and the subsequent convolution module of each group.

[0064] Optionally, step S12 includes at least one of the following:

[0065] Determine or obtain at least one first intermediate value based on the neural network and at least one input feature, and perform filtering processing on at least one image block based on the at least one first intermediate value and at least one lookup table;

[0066] Determine or obtain at least one second intermediate value based on at least one input feature and at least one lookup table, and perform filtering on at least one image block based on the neural network and the at least one second intermediate value;

[0067] Determine or generate at least one index according to at least one input feature, and perform filtering processing on at least one image block according to the at least one index and at least one lookup table;

[0068] Determining or obtaining at least one third intermediate value based on the neural network and the at least one input feature, determining or obtaining at least one fourth intermediate value based on the at least one third intermediate value and at least one lookup table, and filtering the at least one image block based on the neural network and the at least one fourth intermediate value;

[0069] At least one fifth intermediate value is determined or obtained based on at least one input feature and at least one lookup table, at least one sixth intermediate value is determined or obtained based on the at least one fifth intermediate value and a neural network, and filtering is performed on at least one image block based on the at least one sixth intermediate value and the at least one lookup table.

[0070] Optionally, the derivative block is determined or obtained by at least one of the following:

[0071] a cropping result of cropping a reconstructed block and / or a predicted block of at least one image block;

[0072] a filling result of filling a reconstructed block and / or a predicted block of at least one image block;

[0073] an update result of performing pixel update on a reconstructed block and / or a predicted block of at least one image block;

[0074] a translation result of performing pixel translation on a reconstructed block and / or a predicted block of at least one image block;

[0075] An editing result of performing pixel editing on a reconstructed block and / or a prediction block of at least one image block.

[0076] The present application also provides a processing device, comprising:

[0077] The processing module is used to perform filtering processing on at least one image block based on multiple input branches, a neural network and / or a lookup table.

[0078] The present application also provides a processing device, comprising: a memory and a processor, wherein an image processing program is stored in the memory, and when the image processing program is executed by the processor, the steps of any of the above-mentioned image processing methods are implemented.

[0079] The present application also provides a storage medium storing a computer program, wherein the computer program, when executed by a processor, implements the steps of any of the above-mentioned image processing methods.

[0080] As described above, the image processing method of the present application can be applied to a processing device, including: filtering at least one image block based on multiple input branches, a neural network, and / or a lookup table. Through the technical solution of the present application, when filtering at least one image block using a neural network and / or a lookup table, multiple input branches are comprehensively considered, which can reduce the complexity of the filtering process and thereby support improved efficiency of video encoding and / or decoding. BRIEF DESCRIPTION OF THE DRAWINGS

[0081] The accompanying drawings herein are incorporated into and constitute a part of the specification, illustrate embodiments consistent with the present application, and together with the specification, are used to explain the principles of the present application. In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following is a brief introduction to the drawings required for describing the embodiments. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without inventive work.

[0082] Figure 1 A schematic diagram of the hardware structure of a mobile terminal for implementing various embodiments of the present application;

[0083] Figure 2 A communication network system architecture diagram provided in an embodiment of the present application;

[0084] Figure 3 A schematic diagram of the hardware structure of a controller 140 provided in this application;

[0085] Figure 4 A schematic diagram of the hardware structure of a network node 150 provided in this application;

[0086] Figure 5 A DCT transform diagram provided by this application;

[0087] Figure 6 is a flowchart of an image processing method according to the first embodiment;

[0088] Figure 7 A schematic diagram of the encoding and decoding process in the image processing method provided in this application;

[0089] Figure 8 is a flowchart of an image processing method according to a second embodiment;

[0090] Figure 9 A schematic diagram of the architecture for processing based on a neural network and a first input branch provided in this application;

[0091] Figure 10 It is a schematic diagram of a processing module of a processing device.

[0092] The purpose of this application, its features, and advantages will be further described in conjunction with the embodiments and with reference to the accompanying drawings. The accompanying drawings illustrate specific embodiments of this application, which will be described in more detail below. These drawings and the accompanying text are not intended to limit the scope of the present application in any way, but rather to illustrate the concepts of this application to those skilled in the art by reference to specific embodiments. DETAILED DESCRIPTION

[0093] Exemplary embodiments will be described in detail herein, with examples illustrated in the accompanying drawings. In the following description, when referring to the drawings, identical numerals in different figures represent identical or similar elements, unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all embodiments consistent with the present application. Rather, they are merely examples of apparatus and methods consistent with certain aspects of the present application, as detailed in the appended claims.

[0094] It should be noted that, in this document, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the sentence "comprising a ..." does not exclude the presence of other identical elements in the process, method, article or device comprising the element. In addition, components, features, and elements with the same name in different embodiments of the present application may have the same meaning or different meanings, and their specific meanings need to be determined by their explanation in the specific embodiment or further combined with the context of the specific embodiment.

[0095] It should be understood that although the terms first, second, third, etc. may be used herein to describe various information, such information should not be limited to these terms. These terms are used only to distinguish information of the same type from one another. For example, without departing from the scope of this document, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the term "if" as used herein may be interpreted as "when," "when," or "in response to a determination." Furthermore, as used herein, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context indicates otherwise. It should be further understood that the terms "comprising" and "including" indicate the presence of the described features, steps, operations, elements, components, items, types, and / or groups, but do not exclude the presence, occurrence, or addition of one or more other features, steps, operations, elements, components, items, types, and / or groups. The terms "or," "and / or," "including at least one of the following," etc., used herein, may be interpreted as inclusive, or mean any one or any combination. For example, “comprising at least one of the following: A, B, C” means “any of the following: A; B; C; A and B; A and C; B and C; A and B and C”; and for another example, “A, B or C” or “A, B and / or C” means “any of the following: A; B; C; A and B; A and C; B and C; A and B and C”. An exception to this definition will occur only when a combination of elements, functions, steps or operations are inherently mutually exclusive in some manner.

[0096] It should be understood that, although the various steps in the flowchart in the embodiment of the present application are shown in sequence according to the indication of the arrows, these steps are not necessarily performed in sequence in the order indicated by the arrows. Unless clearly stated herein, the execution of these steps is not strictly limited in order, and they can be performed in other orders. Moreover, at least a portion of the steps in the figure may include multiple sub-steps or multiple stages, and these sub-steps or stages are not necessarily performed at the same time, but can be performed at different times, and their execution order is not necessarily performed in sequence, but can be performed in turn or alternately with at least a portion of other steps or sub-steps or stages of other steps.

[0097] As used herein, the words "if" and "if" may be interpreted as "at the time of" or "when" or "in response to determining" or "in response to detecting," depending on the context. Similarly, the phrases "if it is determined" or "if (stated condition or event) is detected" may be interpreted as "when it is determined" or "in response to the determination" or "when detecting (stated condition or event)" or "in response to detecting (stated condition or event)," depending on the context.

[0098] It should be noted that in this article, step codes such as S10 and S20 are used for the purpose of expressing the corresponding content more clearly and concisely, and do not constitute a substantial limitation on the order. When implementing the step, those skilled in the art may execute S20 first and then S10, etc., but these should all be within the scope of protection of this application.

[0099] It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0100] In the subsequent description, the use of suffixes such as "module", "component" or "unit" to represent elements is only for the purpose of facilitating the description of the present application and has no specific meaning. Therefore, "module", "component" or "unit" can be used interchangeably.

[0101] The processing device in this application can be a smart terminal or a server, etc., and the smart terminal can be implemented in various forms. For example, the smart terminal described in this application can include smart terminals such as mobile phones, tablet computers, laptop computers, PDAs, portable media players (PMPs), navigation devices, wearable devices, smart bracelets, pedometers, etc., as well as fixed terminals such as digital TVs and desktop computers.

[0102] The subsequent description will be made by taking a mobile terminal as an example. It will be understood by those skilled in the art that, in addition to components specifically used for mobile purposes, the configuration according to the embodiments of the present application can also be applied to fixed-type terminals.

[0103] See also Figure 1 , which is a schematic diagram of the hardware structure of a mobile terminal for implementing various embodiments of the present application. The mobile terminal 100 may include: an RF (Radio Frequency) unit 101, a WiFi module 102, an audio output unit 103, an A / V (audio / video) input unit 104, a sensor 105, a display unit 106, a user input unit 107, an interface unit 108, a memory 109, a processor 110, and a power supply 111. Those skilled in the art will understand that Figure 1 The structure of the mobile terminal shown in the figure does not constitute a limitation to the mobile terminal. The mobile terminal may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently.

[0104] The following combination Figure 1 A detailed introduction to the various components of the mobile terminal:

[0105] The RF unit 101 can be used to send and receive information or receive signals during calls. Specifically, it receives downlink information from the base station and transmits it to the processor 110 for processing. It also transmits uplink data to the base station. Typically, the RF unit 101 includes, but is not limited to, an antenna, at least one amplifier, a transceiver, a coupler, a low-noise amplifier, a duplexer, and more. Furthermore, the RF unit 101 can communicate with the network and other devices via wireless communication. The above-mentioned wireless communications may use any communication standard or protocol, including but not limited to GSM (Global System of Mobile communication), GPRS (General Packet Radio Service), CDMA2000 (Code Division Multiple Access 2000), WCDMA (Wideband Code Division Multiple Access), TD-SCDMA (Time Division-Synchronous Code Division Multiple Access), FDD-LTE (Frequency Division Duplexing-Long Term Evolution), TDD-LTE (Time Division Duplexing-Long Term Evolution), 5G and 6G, etc.

[0106] WiFi is a short-range wireless transmission technology. Mobile terminals can help users send and receive emails, browse web pages, and access streaming media through the WiFi module 102. It provides users with wireless broadband Internet access. Figure 1 The WiFi module 102 is shown, but it is understandable that it is not an essential component of the mobile terminal and can be omitted as needed without changing the essence of the invention.

[0107] The audio output unit 103 can convert audio data received by the RF unit 101 or the WiFi module 102 or stored in the memory 109 into an audio signal and output it as sound when the mobile terminal 100 is in a call signal reception mode, a talk mode, a recording mode, a voice recognition mode, a broadcast reception mode, or the like. Furthermore, the audio output unit 103 can also provide audio output related to a specific function performed by the mobile terminal 100 (e.g., a call signal reception sound, a message reception sound, etc.). The audio output unit 103 may include a speaker, a buzzer, or the like.

[0108] The A / V input unit 104 is used to receive audio or video signals. The A / V input unit 104 may include a graphics processing unit (GPU) 1041 and a microphone 1042. The GPU 1041 processes image data of still images or videos captured by an image capture device (e.g., a camera) in video capture mode or image capture mode. The processed image frames may be displayed on the display unit 106. The image frames processed by the GPU 1041 may be stored in the memory 109 (or other storage medium) or transmitted via the RF unit 101 or the WiFi module 102. The microphone 1042 may receive sound (audio data) in operating modes such as a phone call mode, a recording mode, and a voice recognition mode, and may process such sound into audio data. In the phone call mode, the processed audio (voice) data may be converted into a format that can be transmitted to a mobile communication base station via the RF unit 101. The microphone 1042 may implement various types of noise cancellation (or suppression) algorithms to eliminate (or suppress) noise or interference generated during the reception and transmission of audio signals.

[0109] The mobile terminal 100 also includes at least one sensor 105, such as a light sensor, a motion sensor, and other sensors. Optionally, the light sensor includes an ambient light sensor and a proximity sensor. Optionally, the ambient light sensor can adjust the brightness of the display panel 1061 according to the brightness of the ambient light, and the proximity sensor can turn off the display panel 1061 and / or the backlight when the mobile terminal 100 is moved to the ear. As a type of motion sensor, the accelerometer sensor can detect the magnitude of acceleration in all directions (generally three axes), and can detect the magnitude and direction of gravity when stationary. It can be used for applications that recognize the posture of the mobile phone (such as horizontal and vertical screen switching, related games, magnetometer posture calibration), vibration recognition related functions (such as pedometer, tapping), etc.; as for other sensors that can be configured in the mobile phone, such as fingerprint sensors, pressure sensors, iris sensors, molecular sensors, gyroscopes, barometers, hygrometers, thermometers, infrared sensors, etc., they will not be described here.

[0110] The display unit 106 is used to display information input by the user or information provided to the user. The display unit 106 may include a display panel 1061, which may be configured in the form of a liquid crystal display (LCD), an organic light-emitting diode (OLED), or the like.

[0111] The user input unit 107 can be used to receive input digital or character information, and to generate key signal input related to the user settings and function control of the mobile terminal. Optionally, the user input unit 107 may include a touch panel 1071 and other input devices 1072. The touch panel 1071, also known as a touch screen, can collect user touch operations on or near it (such as operations performed by the user using a finger, stylus, or any other suitable object or accessory on or near the touch panel 1071) and drive the corresponding connection device according to a pre-set program. The touch panel 1071 may include two parts: a touch detection device and a touch controller. Optionally, the touch detection device detects the user's touch direction and detects the signal generated by the touch operation, and transmits the signal to the touch controller; the touch controller receives the touch information from the touch detection device and converts it into touch point coordinates, which are then sent to the processor 110. It can also receive commands sent by the processor 110 and execute them. In addition, the touch panel 1071 can be implemented using various types such as resistive, capacitive, infrared, and surface acoustic wave. In addition to the touch panel 1071, the user input unit 107 may further include other input devices 1072. Optionally, the other input devices 1072 may include, but are not limited to, one or more of a physical keyboard, function keys (such as volume control keys, power keys, etc.), a trackball, a mouse, a joystick, etc., and the specifics are not limited here.

[0112] Optionally, the touch panel 1071 may cover the display panel 1061. When the touch panel 1071 detects a touch operation on or near it, it transmits the information to the processor 110 to determine the type of touch event. The processor 110 then provides a corresponding visual output on the display panel 1061 according to the type of touch event. Figure 1 In the embodiment, the touch panel 1071 and the display panel 1061 are two independent components to realize the input and output functions of the mobile terminal. However, in some embodiments, the touch panel 1071 and the display panel 1061 can be integrated to realize the input and output functions of the mobile terminal, which is not limited here.

[0113] The interface unit 108 serves as an interface through which at least one external device can be connected to the mobile terminal 100. For example, the external device may include a wired or wireless headset port, an external power supply (or battery charger) port, a wired or wireless data port, a memory card port, a port for connecting a device with an identification module, an audio input / output (I / O) port, a video I / O port, a headphone port, etc. The interface unit 108 may be used to receive input (e.g., data information, power, etc.) from an external device and transmit the received input to one or more elements within the mobile terminal 100 or may be used to transmit data between the mobile terminal 100 and an external device.

[0114] Memory 109 can be used to store software programs and various data. Memory 109 may primarily include a program storage area and a data storage area. Optionally, the program storage area may store an operating system and at least one application required for a function (such as a sound playback function or an image playback function); the data storage area may store data generated based on the use of the mobile phone (such as audio data, a phone book, etc.). Furthermore, memory 109 may include high-speed random access memory and non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device.

[0115] Processor 110 is the control center of the mobile terminal, connecting all components of the mobile terminal using various interfaces and circuits. By running or executing software programs and / or modules stored in memory 109 and accessing data stored in memory 109, it executes various functions of the mobile terminal and processes data, thereby providing overall monitoring of the mobile terminal. Processor 110 may include one or more processing units; preferably, processor 110 may integrate an application processor and a modem processor. Optionally, the application processor primarily handles the operating system, user interface, and application programs, while the modem processor primarily handles wireless communications. It is understood that the modem processor may not be integrated into processor 110.

[0116] The mobile terminal 100 may also include a power supply 111 (such as a battery) for supplying power to various components. Preferably, the power supply 111 may be logically connected to the processor 110 through a power management system, thereby managing functions such as charging, discharging, and power consumption through the power management system.

[0117] although Figure 1 Not shown, the mobile terminal 100 may further include a Bluetooth module, etc., which will not be described in detail here.

[0118] To facilitate understanding of the embodiments of the present application, the communication network system on which the mobile terminal of the present application is based is described below.

[0119] See also Figure 2 , Figure 2 A communication network system architecture diagram is provided for an embodiment of the present application. The communication network system is an LTE system of universal mobile communication technology. The LTE system includes a UE (User Equipment) 201, an E-UTRAN (Evolved UMTS Terrestrial Radio Access Network) 202, an EPC (Evolved Packet Core) 203 and an operator's IP service 204, which are connected in sequence.

[0120] Optionally, UE201 may be the above-mentioned terminal 100, which will not be described in detail here.

[0121] E-UTRAN 202 includes eNodeB 2021 and other eNodeBs 2022 . Optionally, eNodeB 2021 may be connected to other eNodeBs 2022 via a backhaul (eg, an X2 interface). eNodeB 2021 is connected to EPC 203 , and eNodeB 2021 may provide access from UE 201 to EPC 203 .

[0122] EPC 203 may include an MME (Mobility Management Entity) 2031, an HSS (Home Subscriber Server) 2032, other MMEs 2033, an SGW (Serving Gate Way) 2034, a PGW (PDN Gate Way) 2035, and a PCRF (Policy and Charging Rules Function) 2036. Optionally, MME 2031 is a control node that processes signaling between UE 201 and EPC 203, providing bearer and connection management. HSS 2032 provides registers for managing functions such as the Home Location Register (not shown) and stores user-specific information such as service features and data rates. All user data can be sent through SGW2034, PGW2035 can provide IP address allocation and other functions for UE 201, PCRF2036 is the policy and charging control policy decision point for service data flow and IP bearer resources, and it selects and provides available policy and charging control decisions for the policy and charging execution function unit (not shown in the figure).

[0123] The IP service 204 may include the Internet, an intranet, an IMS (IP Multimedia Subsystem), or other IP services.

[0124] Although the above introduction takes the LTE system as an example, those skilled in the art should know that this application is not only applicable to the LTE system, but can also be applied to other wireless communication systems, such as GSM, CDMA2000, WCDMA, TD-SCDMA, 5G and future new network systems (such as 6G), etc., which are not limited here.

[0125] Figure 3This is a schematic diagram of the hardware structure of a controller 140 provided in this application. The controller 140 includes a memory 1401 and a processor 1402. The memory 1401 is used to store program instructions, and the processor 1402 is used to call the program instructions in the memory 1401 to execute the steps performed by the controller in the first embodiment of the above method. The implementation principles and beneficial effects are similar and will not be repeated here.

[0126] Optionally, the controller further includes a communication interface 1403, which can be connected to the processor 1402 via a bus 1404. The processor 1402 can control the communication interface 1403 to implement the receiving and sending functions of the controller 140.

[0127] Figure 4 This is a schematic diagram of the hardware structure of a network node 150 provided in this application. The network node 150 includes a memory 1501 and a processor 1502. The memory 1501 is used to store program instructions, and the processor 1502 is used to call the program instructions in the memory 1501 to execute the steps performed by the first node in the first embodiment of the above method. The implementation principles and beneficial effects are similar and will not be repeated here.

[0128] Optionally, the controller further includes a communication interface 1503, which can be connected to the processor 1502 via a bus 1504. The processor 1502 can control the communication interface 1503 to implement the receiving and sending functions of the network node 150.

[0129] The integrated modules implemented in the form of software function modules can be stored in a computer-readable storage medium. The software function modules stored in a storage medium include a number of instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) or a processor to execute some of the steps of the methods of various embodiments of the present application.

[0130] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When software is used for implementation, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function according to the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a storage medium or transmitted from one storage medium to another storage medium. For example, the computer instructions can be transmitted from a website, computer, server or data center to another website, computer, server or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) method. The storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more available media integrations. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid state drive solid state disk, SSD), etc.

[0131] Based on the above-mentioned mobile terminal hardware structure and communication network system, various embodiments of the present application are proposed.

[0132] Optionally, a brief introduction is given to the neural networks that may be involved in the embodiments of the present application.

[0133] The NNLF (Neural Network based Loop Filter) structure in NNVC (Neural Network based Video Coding) includes two structures: HOP (High Complexity Operation Point) and LOP (Low Complexity Operation Point). For the LOP structure, the input of the loop filter may include reconstructed luma and chroma samples (Rec), predicted luma and chroma samples (Pred), luma and chroma boundary strength information (BS), base quantization parameter (QPbase), slice quantization parameter (QPslice), and block prediction information (IPB).

[0134] The luma samples input by Rec and Pred are transformed. For each W (width) × H (height) input luma block, a 2 × 2 DCT-II transform is applied to each 2 × 2 sub-block. The result is reconstructed into a (W / 2) × (H / 2) × 4 tensor, where 4 represents the four frequency channels. The transformed luma samples are then concatenated with the chroma U and V channels to form a feature tensor of (W / 2) × (H / 2) × 6.

[0135] Optionally, when constructing the reconstructed sample example, 8 neighboring samples can be expanded in each direction to create a 144×144 image block. This 144×144×1 input luminance block is transformed and reconstructed into a 72×72×4 tensor. For the chrominance reconstructed samples, since the values within each 2×2 sub-block remain constant due to the 2x upsampling, no transformation is required. It can be directly downsampled by 2 to obtain a 72×72×2 tensor. To maintain compatibility with the current training process, the 2x upsampling operation of chrominance is retained because in the training data, the chrominance data has been oversampled and saved together with the luminance data. In a more efficient implementation, both upsampling and downsampling operations can be omitted without affecting the final result.

[0136] For QPbase (base quantization parameter) and QPsl ice (slice quantization parameter), since they remain constant within the entire image block, no transformation is required and they are directly reconstructed into a tensor of (W / 2) × (H / 2) × 1. For IPB (block prediction information), since it only contains brightness information, it is transformed and reconstructed into a tensor of (W / 2) × (H / 2) × 4.

[0137] The transformed input data is processed through a convolutional layer with kernel sizes of 3×3 or 1×1, followed by concatenation for feature fusion and transformation. In the fusion and transformation modules, 1×3 and 3×1 separable convolutions and a downsampling operation of a factor of 2 are used. The network then splits into two branches, one for processing luminance information and the other for processing chrominance information.

[0138] It can be seen that the NNLF based on the 2x2DCT transform only processes independent non-overlapping image blocks during the image feature extraction process, resulting in it completely ignoring the spatial correlation information of the overlapping pixel areas at the boundaries of adjacent DCT (Discrete Cosine Transform) blocks. In other words, the image block is divided into non-overlapping 2x2 sub-blocks, and each sub-block is independently DCT transformed. The cross-block correlation of pixels at the boundary of adjacent sub-blocks (such as edge continuity and texture consistency) is not considered, that is, the overlapping area information is not considered. Accordingly, this type of information can be called overlapping area information or sub-block boundary information.

[0139] like Figure 5 As shown in the figure, the DCT transform only considers the information in the solid box in the image block, but ignores the information in the dotted box. This lack of information makes it impossible for the filter to effectively capture the transition features between image blocks, resulting in the loss of key structural details during the feature extraction process, which in turn affects the subsequent filtering operation's ability to smooth the image block boundaries and the overall reconstruction quality.

[0140] Optionally, since the traditional NNLF uses 2x2DCT transform to extract image features of image blocks, the performance is limited due to ignoring the overlapping area information of the image blocks. Therefore, in an embodiment of the present application, the image blocks can be cropped and filled to obtain a derivative block of at least one image block, so that the derivative block fully integrates the spatial correlation of the boundary pixels of adjacent DCT blocks. The derivative block can contain the overlapping area information or sub-block boundary information of the image block, thereby improving the filtering performance while maintaining low complexity. And / or, for the defect that when the Y channel and UV channel in the original NNLF are processed independently, the newly added input branch will lead to an increase in network complexity. In this embodiment, corresponding optimized channel processing logic will also be set, such as designing a network architecture that fuses the Y component with the new Y component and the UV component with the new UV component in parallel, to effectively reduce computational redundancy and achieve balanced optimization of model complexity and performance.

[0141] Optionally, in an embodiment of the present application, a cross-block modeling and multi-directional extraction mechanism of overlapping pixel features can be implemented.

[0142] It can break through the limitations of the unprocessed sub-block boundary information of 2X2 DCT, and / or can capture the overlapping area information of the adjacent block boundaries in the horizontal, vertical, and horizontal and vertical mixed directions through operations such as cropping and padding, so that the filter can perceive the spatial correlation of cross-block pixels (such as edge continuity and texture transition features), avoid the defect of incomplete feature extraction caused by block segmentation, and by designing a multi-directional overlapping information extraction process, through lightweight operations such as zero padding and mirror symmetric padding, expand the perception range of DCT transform without almost increasing parameters, and improve the spatial context integrity of features.

[0143] Optionally, in an embodiment of the present application, a low-complexity multi-component fusion architecture based on grouped convolution can be constructed.

[0144] Optionally, to avoid the drawback of increased computational complexity caused by the addition of overlapping area information or sub-block boundary information, in an embodiment of the present application, the Y luminance component (original Y + overlapped Y) and the UV chrominance component (original UV + overlapped UV) can be integrated into independent branches to avoid redundant computations in traditional independent channel processing. A grouped convolution with a grouping coefficient of 2 is introduced in the feature fusion stage, dividing the input channels into two groups for parallel processing. While maintaining the interaction of multi-component features, the computational complexity is reduced by approximately 50%, achieving a balanced optimization of complexity and performance.

[0145] Optionally, in an embodiment of the present application, a multi-branch collaborative channel feature joint optimization mechanism can be implemented.

[0146] In the Y branch, after channel-wise concatenation of the DCT features of the original and overlapping blocks, feature fusion is achieved through group convolution, enhancing edge structure information dominated by brightness. In the UV branch, a similar process is used to preserve color transition details in chroma, while further reducing overall complexity through cross-branch parameter sharing. This overcomes the limitations of the traditional LOP model, which requires independent processing of the Y and UV channels, and achieves collaborative optimization of brightness and chroma features in overlapping regions through component concatenation and group convolution, improving the overall consistency of image reconstruction.

[0147] First embodiment

[0148] Reference Figure 6 , Figure 6 FIG. 1 is a flow chart of an image processing method according to a first embodiment. The image processing method according to the embodiment of the present application can be applied to a processing device, including step S10:

[0149] S10 , performing filtering processing on at least one image block according to the multiple input branches, the neural network and / or the lookup table.

[0150] In this embodiment, the processing device can be a smart terminal, such as a mobile phone, a computer, etc., or a server, such as a local server or a cloud server. In this embodiment and this application, the processing device is mainly described as a smart terminal.

[0151] Optionally, the technical solution of this embodiment can be applied to the fields of image coding and decoding, video coding and decoding, hardware video coding and decoding, dedicated circuit video coding and decoding, real-time video coding and decoding, and so on.

[0152] Optionally, for ease of understanding, a brief introduction to the encoding and decoding process is given: Figure 7As shown, it includes modules such as general coding control, transformation and quantization, intra-frame estimation, intra-frame prediction, motion compensation, motion estimation, inverse quantization and inverse transformation, filter control analysis, deblocking filter and SAO filter (i.e., loop filter), entropy coding, and decoding frame buffer. Optionally, the motion compensation module can perform intra-frame / inter-frame selection to determine specific compensation. Optionally, when performing entropy coding, it is based on the general control data determined by the general coding control module, the variable quantization coefficient determined by the transformation and quantization module, the intra-frame prediction data determined by the filter control analysis, and the motion data determined by the filter control and decoding frame buffer, thereby obtaining the coding bit rate.

[0153] Optionally, the decoded video signal is outputted via a decoded frame buffer.

[0154] Optionally, the loop filter may include two branches: a deblocking filter and an LC-NNLF. The branch results are then fused and SAO (Sample Adaptive Offset) and ALF (Adaptive Loop Filter) processing is performed.

[0155] Optionally, the multiple input branches can be different input branches for inputting into a neural network and / or a lookup table. Multiple input branches can be constructed according to the luminance component and the chrominance component, and the luminance information and chrominance information corresponding to at least one image block can be input into the neural network and / or the lookup table according to the respective corresponding input branches for filtering processing. Other rules can also be used, such as comparing the reference pixel value of at least one image block with a pixel threshold, selecting an input branch for reference pixel values greater than the pixel threshold to input into the neural network and / or the lookup table, and selecting another input branch for reference pixel values less than or equal to the pixel threshold to input into the neural network and / or the lookup table for filtering processing, etc.

[0156] Optionally, the neural network can be a neural network based on a fully connected layer; a neural network based on a convolutional layer; a neural network based on a Transformer; a neural network based on a hybrid convolutional layer, a fully connected layer, and a Transformer, etc. In this embodiment, only a neural network based on a convolutional layer is used as an example.

[0157] Optionally, the lookup table may be a filtering lookup table, which may include a filtering mode, and / or a filtering pixel value, etc.

[0158] Optionally, the input parameters in each input branch within the multiple input branches can be determined based on at least one image block, or based on a reconstructed block and / or a predicted block of at least one image block, or based on a derivative block of at least one reconstructed block and / or a predicted block.

[0159] Optionally, the derivative block may be determined or obtained based on a cropping result of cropping the prediction block and / or the reconstructed block of at least one image block.

[0160] Optionally, the derivative block may be determined or obtained based on a filling result of filling the prediction block and / or the reconstructed block of at least one image block.

[0161] Optionally, the derivative block may be determined or obtained based on the result of pixel editing, pixel shifting, and / or pixel updating of the prediction block and / or the reconstructed block of at least one image block.

[0162] Optionally, when filtering at least one image block, a derivative block of the at least one image block can be comprehensively considered, and then when filtering at least one image block through the derivative block, spatial correlation information of overlapping pixel areas at the boundaries of adjacent DCT blocks can be comprehensively considered.

[0163] Optionally, a transformation process may be performed on a derivative block of at least one image block, such as performing a DCT operation on the derivative block to extract derivative block information. The derivative block information may include RecOverlap (reconstructed block overlap region information) and PredOverlap (prediction block overlap region information), reconstructed block sub-block boundary information, prediction block sub-block boundary information, etc. The derivative block information may include overlap region information or sub-block boundary information. Derivative block transformation features are determined based on the derived block information, such as by directly using pixel values and / or pixel positions of the derivative block as the derived block transformation features.

[0164] Optionally, transformation processing may be performed on at least one reconstructed block and / or predicted block to obtain corresponding reconstructed transformation features and predicted transformation features.

[0165] Optionally, based on at least one of the reconstructed transformation features, the predicted transformation features, and the derived block transformation features, the features in each input branch of the multiple input branches can be determined and used as input parameters, and then filtered according to a neural network and / or a lookup table to obtain a target image block after filtering.

[0166] Optionally, the processing device may be a decoding end. If at the decoding end, filtering processing may be performed on at least one image block based on multiple input branches, a neural network and / or a lookup table.

[0167] Optionally, the processing device may be an encoding end. If at the encoding end, filtering processing may be performed on at least one image block based on multiple input branches, a neural network and / or a lookup table.

[0168] In an embodiment of the present application, filtering processing is performed on at least one image block based on multiple input branches, a neural network and / or a lookup table, so that when filtering processing is performed on at least one image block using a neural network and / or a lookup table, multiple input branches are comprehensively considered, which can reduce the complexity of the filtering processing and thereby support improving the efficiency of video encoding and / or decoding.

[0169] Second embodiment

[0170] Based on the first embodiment, a second embodiment is proposed.

[0171] In this embodiment, the image block feature includes at least one of a reconstruction transformation feature of a reconstructed block of at least one image block, a prediction transformation feature of a prediction block of at least one image block, and a derivative transformation feature of a derivative block of at least one image block.

[0172] Optionally, at least one reconstructed block may be subjected to a transformation process, such as at least one of DCT transform (Discrete Cosine Transform), DST transform (Discrete Sine Transform), KL transform (Karhunen–Loève Transform), wavelet transform, Hadamard transform, etc.

[0173] Optionally, pixel features may be extracted from at least one reconstructed block after the transformation process to obtain reconstructed transformation features, such as pixel positions and pixel values of the reconstructed block.

[0174] Optionally, a transform process may be performed on at least one prediction block.

[0175] Optionally, pixel features may be extracted from at least one prediction block after the transformation process to obtain prediction transformation features, such as pixel positions and pixel values of the prediction block.

[0176] Optionally, the derivative block of at least one image block may include at least one of a derivative block of at least one reconstructed block and a derivative block of at least one prediction block.

[0177] Optionally, transformation processing may be performed on at least one derivative block, and pixel features may be extracted from the at least one derivative block after transformation processing to obtain derivative transformation features, such as pixel positions and pixel values of the derivative block.

[0178] Optionally, the transformation methods for performing the transformation processing on at least two of the reconstructed block, the predicted block and the derived block may be the same or different, and there is no limitation here.

[0179] In this embodiment, by determining that at least one of the reconstructed transformation features, the predicted transformation features, and the derived transformation features is an image block feature, it is convenient to subsequently perform channel splicing based on the image block features in combination with multiple input branches, and then perform filtering processing in combination with a neural network and / or a lookup table.

[0180] Optionally, the channels in this embodiment may be applicable to a lookup table.

[0181] Optionally, a channel is a component of a feature map of an image block in a depth dimension, which is used to describe the feature representation of the number of features in a specific dimension. Each channel represents a certain feature (such as texture, edge, derivative distribution) extracted from an image block (such as a prediction block or a reconstructed block). For example, one channel may be used to detect horizontal edges, and another channel may be used to detect vertical edges.

[0182] Optionally, the image block features of at least two channels can be channel-stitched to obtain a first multi-channel feature, and then based on the correspondence between the channel features and the indexes set in advance, the index corresponding to the first multi-channel feature (such as a one-dimensional index, a two-dimensional index or a three-dimensional index, etc.) is determined or obtained, and input into a lookup table for search to determine or obtain the target image block after filtering.

[0183] Optionally, the derivative block is determined or obtained by at least one of the following steps a1 to a5:

[0184] Step a1, cropping a reconstructed block and / or a predicted block of at least one image block;

[0185] Optionally, for the derivative block of at least one reconstructed block, the boundary area of the at least one reconstructed block (such as at least one item of at least one leftmost column of the reconstructed block, at least one rightmost column of the reconstructed block, at least one topmost row of the reconstructed block, at least one bottommost row of the reconstructed block, etc.) can be cropped to obtain a cropping result.

[0186] Optionally, the cropping result may include at least one cropped reconstructed block.

[0187] Optionally, a derivative block of the at least one reconstructed block may be determined or obtained based on the at least one cropped reconstructed block, for example, by performing zero padding on the at least one cropped reconstructed block.

[0188] Optionally, the size parameters of the derivative block of at least one reconstructed block match the size parameters of at least one reconstructed block, for example, the width of the reconstructed block n is consistent with the width of the derivative block of the reconstructed block n, and the height of the reconstructed block n is consistent with the height of the derivative block of the reconstructed block n.

[0189] Optionally, for a derivative block of at least one prediction block,

[0190] The boundary area of at least one prediction block (such as at least one item of the leftmost column of the prediction block, at least one rightmost column of the prediction block, at least one topmost row of the prediction block, at least one bottommost row of the prediction block, etc.) can be cropped to obtain a cropping result.

[0191] Optionally, the cropping result may include at least one cropped prediction block.

[0192] Optionally, a derivative block of at least one prediction block may be determined or obtained based on the at least one cropped prediction block, for example, by performing mirror symmetric prime filling on the at least one cropped prediction block.

[0193] Optionally, the size parameters of the derivative block of at least one prediction block match the size parameters of at least one prediction block, for example, the width of prediction block m is consistent with the width of the derivative block of prediction block m, and the height of prediction block m is consistent with the height of the derivative block of prediction block m.

[0194] In this embodiment, a derivative block is determined or obtained based on a cropping result of a reconstructed block and / or a predicted block of at least one image block, thereby ensuring that the derivative block is closely associated with the reconstructed block and / or the predicted block. This facilitates subsequent use of the derivative block transformation features of the derivative block in combination with multiple input branches, a neural network and / or a lookup table to perform filtering processing, thereby ensuring the effectiveness of the filtering processing.

[0195] Step a2, filling the reconstructed block and / or the predicted block of at least one image block;

[0196] Optionally, the reconstructed block and / or prediction block of at least one image block may be cropped, and then the cropped at least one reconstructed block and / or prediction block may be filled, and the filling method may include at least one of zero filling and mirror symmetric filling.

[0197] For example, the leftmost column and the rightmost column of the at least one reconstructed block and / or prediction block may be cropped, and then two columns of pixels may be padded at the rightmost side of the at least one reconstructed block and / or prediction block. Alternatively, the topmost row and the bottommost row of the at least one reconstructed block and / or prediction block may be cropped, and then two columns of pixels may be padded at the bottom of the at least one reconstructed block and / or prediction block.

[0198] Optionally, pixel cleaning may be performed on the even columns of the boundary (such as the leftmost and / or rightmost) of at least one reconstructed block and / or predicted block, and then pixel filling may be performed on the area of the boundary of at least one reconstructed block and / or predicted block that has been pixel cleaned, and the filling method may include at least one of zero filling and mirror-symmetric filling.

[0199] Optionally, pixel cleaning may be performed on the even rows at the boundaries (such as the top and / or bottom) of at least one reconstructed block and / or predicted block, and then pixel filling may be performed on the pixel-cleaned region at the boundaries of at least one reconstructed block and / or predicted block, and the filling method may include at least one of zero filling and mirror-symmetric filling.

[0200] Optionally, the filling result includes at least one reconstructed block that has been filled, and the reconstructed block can be used as a derivative block corresponding to the reconstructed block.

[0201] Optionally, the filling result includes at least one padded prediction block, which can be used as a derivative block corresponding to the prediction block.

[0202] In this embodiment, a derivative block is determined or obtained based on a filling result of a reconstructed block and / or a predicted block of at least one image block, thereby ensuring that the derivative block is closely associated with the reconstructed block and / or the predicted block. This facilitates subsequent use of the derivative block transformation features of the derivative block in combination with multiple input branches, a neural network and / or a lookup table to perform filtering processing, thereby ensuring the effectiveness of the filtering processing.

[0203] Step a3, performing pixel updating on the reconstructed block and / or the predicted block of at least one image block;

[0204] Optionally, pixels of even columns at the boundary (such as the leftmost and / or rightmost) of at least one reconstructed block and / or prediction block may be updated, for example, all pixels may be updated to zero.

[0205] Optionally, pixels of even-numbered rows at the boundaries (such as the top and / or bottom) of at least one reconstructed block and / or predicted block may be updated, for example, all pixels may be updated to zero.

[0206] Optionally, the update result includes at least one reconstructed block after pixel update, and the reconstructed block can be used as a derivative block corresponding to the reconstructed block.

[0207] Optionally, the update result includes at least one prediction block after pixel update, which can be used as a derivative block corresponding to the prediction block.

[0208] In this embodiment, a derivative block is determined or obtained based on an update result of pixels being updated on a reconstructed block and / or a predicted block of at least one image block, thereby ensuring that the derivative block is closely associated with the reconstructed block and / or the predicted block. This facilitates subsequent use of the derivative block transformation features of the derivative block in combination with multiple input branches, a neural network and / or a lookup table to perform filtering processing, thereby ensuring the effectiveness of the filtering processing.

[0209] Step a4, performing pixel shifting on the reconstructed block and / or the predicted block of at least one image block to obtain a shift result;

[0210] Optionally, pixel cleaning may be performed on the even columns of the boundary (such as the leftmost and / or rightmost) of at least one reconstructed block and / or predicted block, and then pixel shifting may be performed on the pixel-cleaned region of the boundary of at least one reconstructed block and / or predicted block.

[0211] Optionally, pixel cleaning may be performed on the even rows of the boundary (such as the top and / or bottom) of at least one reconstructed block and / or predicted block, and then pixel shifting may be performed on the pixel-cleaned region of the boundary of at least one reconstructed block and / or predicted block.

[0212] Optionally, the pixel shifting may be to copy pixels in an area of the same size adjacent to the area where the pixel cleaning process has been performed and then shift the pixels to the area where the pixel cleaning process has been performed.

[0213] Optionally, the translation result includes at least one reconstructed block that has undergone pixel translation, and the reconstructed block may be used as a derivative block corresponding to the reconstructed block.

[0214] Optionally, the translation result includes at least one predicted block that has undergone pixel translation, and the predicted block can be used as a derivative block corresponding to the predicted block.

[0215] In this embodiment, a derivative block is determined or obtained based on a translation result of pixels of a reconstructed block and / or a predicted block of at least one image block, thereby ensuring that the derivative block is closely associated with the reconstructed block and / or the predicted block. This facilitates subsequent use of the derivative block transformation features of the derivative block in combination with multiple input branches, a neural network and / or a lookup table to perform filtering processing, thereby ensuring the effectiveness of the filtering processing.

[0216] Step a5: performing pixel editing on the reconstructed block and / or the predicted block of at least one image block.

[0217] Optionally, pixels of even columns at the boundary (such as the leftmost and / or rightmost) of at least one reconstructed block and / or prediction block may be edited, for example, edited to zero pixels.

[0218] Optionally, pixels of even-numbered rows at the boundaries (such as the top and / or bottom) of at least one reconstructed block and / or predicted block may be edited, for example, edited to zero pixels.

[0219] Optionally, the editing result includes at least one reconstructed block that has undergone pixel editing, and the reconstructed block may be used as a derivative block corresponding to the reconstructed block.

[0220] Optionally, the editing result includes at least one prediction block that has undergone pixel editing, and the prediction block can be used as a derivative block corresponding to the prediction block.

[0221] In this embodiment, a derivative block is determined or obtained based on the editing results of pixel editing of a reconstructed block and / or a predicted block of at least one image block, thereby ensuring that the derivative block is closely associated with the reconstructed block and / or the predicted block. This facilitates subsequent use of the derivative block transformation features of the derivative block in combination with multiple input branches, a neural network and / or a lookup table to perform filtering processing, thereby ensuring the effectiveness of the filtering processing.

[0222] In this embodiment, referring to Figure 8 , step S1 includes step S11 and step S12:

[0223] Step S11, performing channel splicing based on multiple input branches and at least one image block feature to determine or obtain input features;

[0224] Optionally, at least one image block corresponds to multiple channels, and each channel may correspond to at least one image block feature.

[0225] Optionally, each input branch in the multiple input branches may include image block features of at least one channel, and the at least one image block feature in the multiple input branches may be channel-joined according to the channel dimension to obtain a multi-channel feature, which is used as the input feature.

[0226] Optionally, the multi-input branch includes at least one of the following modes:

[0227] Mode 1, for processing the first input branch of the luminance component;

[0228] Optionally, various input branches for channel splicing of different components may be set in advance.

[0229] Optionally, in the first input branch, channel splicing processing can be performed on the brightness component of the image block, such as performing channel splicing on the first feature of the derivative block on the brightness component and the second feature of the reconstructed block and / or the predicted block on the brightness component, and then determining or obtaining the input of the neural network and / or lookup table based on the result of the channel splicing.

[0230] In this embodiment, by determining or obtaining the input of the neural network and / or the lookup table based on the first input branch for processing the brightness component, all information of the brightness component can be integrated together, and then filtered through the neural network and / or the lookup table, thereby ensuring the effectiveness of the filtering process.

[0231] Optionally, in mode 1, the first input branch includes at least one of modes 1 to 5:

[0232] Mode 1, a first sub-input branch for processing the reconstructed transformation features and the derived transformation features on the luminance component;

[0233] Optionally, the first input branch includes a first sub-input branch, and the reconstructed transformation features of the reconstructed block on the luminance component and the derived transformation features of the derivative block corresponding to the reconstructed block can be processed in the first sub-input branch, such as channel-splicing the reconstructed transformation features and the derived transformation features to obtain multi-channel features, which can be used as input features to be input into a lookup table and / or neural network for filtering processing.

[0234] In this embodiment, by determining or obtaining the input of the neural network and / or the lookup table based on the first sub-input branch, the reconstructed transformation characteristics of the reconstructed block of the luminance component and the derived transformation characteristics of the derivative block can be integrated together, and then filtering processing is performed through the neural network and / or the lookup table, thereby ensuring the effectiveness of the filtering processing.

[0235] Mode 2, a second sub-input branch for processing the predicted transformation features and the derived transformation features on the luminance component;

[0236] Optionally, the first input branch includes a second sub-input branch, and the predicted transformation features of the prediction block on the luminance component and the derived transformation features of the derivative block corresponding to the prediction block can be processed in the second sub-input branch, such as channel-splicing the predicted transformation features and the derived transformation features to obtain multi-channel features, which can be used as input features to be input into a lookup table and / or neural network for filtering processing.

[0237] In this embodiment, by determining or obtaining the input of the neural network and / or the lookup table based on the second sub-input branch, the predicted transformation characteristics of the prediction block of the luminance component and the derived transformation characteristics of the derivative block can be integrated together, and then filtering processing is performed through the neural network and / or the lookup table, thereby ensuring the effectiveness of the filtering processing.

[0238] Mode 3, a third sub-input branch for processing the reconstructed transformation features on the luminance component;

[0239] Optionally, the first input branch includes a third sub-input branch, and the reconstruction transformation features of the reconstructed block can be processed in the third sub-input branch. For example, channel splicing can be performed on all reconstruction transformation features on the brightness component of at least one reconstructed block. Other processing operations can also be performed, such as converting them into a format that can be recognized by the neural network and / or lookup table, etc., to obtain corresponding input features so that they can be input into the lookup table and / or neural network for filtering processing.

[0240] In this embodiment, by determining or obtaining the input of the neural network and / or the lookup table based on the third sub-input branch, the reconstruction transformation features of all channels of the reconstructed block of the luminance component can be integrated together, and then filtering processing is performed through the neural network and / or the lookup table, thereby ensuring the effectiveness of the filtering processing.

[0241] Mode 4, a fourth sub-input branch for processing the prediction transformation feature on the luminance component;

[0242] Optionally, the first input branch includes a fourth sub-input branch, and the predicted transformation features of the prediction block can be processed in the fourth sub-input branch. For example, channel splicing can be performed on all predicted transformation features on the brightness component of at least one prediction block. Other processing operations can also be performed, such as converting them into a format that can be recognized by the neural network and / or lookup table, etc., to obtain corresponding input features so that they can be input into the lookup table and / or neural network for filtering processing.

[0243] In this embodiment, by determining or obtaining the input of the neural network and / or the lookup table based on the fourth sub-input branch, the predicted transformation features of all channels of the prediction block of the luminance component can be integrated together, and then filtered through the neural network and / or the lookup table, thereby ensuring the effectiveness of the filtering processing.

[0244] Mode 5 is a fifth sub-input branch for processing the derived transformation feature on the luminance component.

[0245] Optionally, the first input branch includes a fifth sub-input branch, and the derivative transformation features corresponding to the derivative blocks of the prediction block and / or the reconstructed block can be processed in the fifth sub-input branch. For example, channel splicing can be performed on all derivative transformation features on the brightness component of at least one derivative block, and other processing operations can also be performed, such as converting them into a format that can be recognized by the neural network and / or lookup table, etc., to obtain the corresponding input features so that they can be input into the lookup table and / or neural network for filtering processing.

[0246] In this embodiment, by determining or obtaining the input of the neural network and / or the lookup table based on the fifth sub-input branch, it is possible to integrate the derivative transformation features of all channels of the derivative block of the luminance component, and then perform filtering processing through the neural network and / or the lookup table, thereby ensuring the effectiveness of the filtering processing.

[0247] Mode 2, for processing the second input branch of the chrominance component;

[0248] Optionally, in the second input branch, channel splicing processing can be performed on the chrominance components of the image block, such as performing channel splicing on the first feature of the derivative block on the chrominance component and the second feature of the reconstructed block and / or the predicted block on the chrominance component, and then determining or obtaining the input of the neural network and / or lookup table based on the result of the channel splicing.

[0249] In this embodiment, by determining or obtaining the input of the neural network and / or the lookup table based on the second input branch for processing the chrominance component, all information of the chrominance component can be integrated together, and then filtered through the neural network and / or the lookup table, thereby ensuring the effectiveness of the filtering process.

[0250] Optionally, in the second mode, the second input branch includes at least one of modes 6 to 12:

[0251] Mode 6, a sixth sub-input branch for processing the reconstructed transformation features and the derived transformation features on the chrominance component;

[0252] Optionally, the second input branch includes a sixth sub-input branch, and the reconstruction transformation features of the reconstructed block on the chrominance component and the derivative transformation features of the derivative block corresponding to the reconstructed block can be processed in the sixth sub-input branch, such as channel-splicing the reconstruction transformation features and the derivative transformation features to obtain multi-channel features, which can be used as input features to be input into a lookup table and / or neural network for filtering processing.

[0253] In this embodiment, by determining or obtaining the input of the neural network and / or the lookup table based on the sixth sub-input branch, the reconstructed transformation characteristics of the reconstructed block of the chrominance component and the derived transformation characteristics of the derivative block can be integrated together, and then filtering processing is performed through the neural network and / or the lookup table, thereby ensuring the effectiveness of the filtering processing.

[0254] Mode 7, a seventh sub-input branch for processing the predicted transform feature and the derived transform feature on the chrominance component;

[0255] Optionally, the second input branch includes a seventh sub-input branch, and the predicted transformation features of the prediction block on the luminance component and the derived transformation features of the derivative block corresponding to the prediction block can be processed in the seventh sub-input branch, such as channel-splicing the predicted transformation features and the derived transformation features to obtain multi-channel features, which can be used as input features to be input into a lookup table and / or a neural network for filtering processing.

[0256] In this embodiment, by determining or obtaining the input of the neural network and / or the lookup table based on the seventh sub-input branch, the predicted transformation characteristics of the prediction block of the chrominance component and the derived transformation characteristics of the derivative block can be integrated together, and then filtering processing is performed through the neural network and / or the lookup table, thereby ensuring the effectiveness of the filtering processing.

[0257] Mode 8, an eighth sub-input branch for processing the reconstructed transformation feature on the chrominance component;

[0258] Optionally, the second input branch includes an eighth sub-input branch, and the reconstruction transformation features of the reconstructed block can be processed in the eighth sub-input branch. For example, channel splicing can be performed on all reconstruction transformation features on the chrominance component of at least one reconstructed block, and other processing operations can also be performed, such as converting them into a format that can be recognized by the neural network and / or lookup table, etc., to obtain corresponding input features so that they can be input into the lookup table and / or neural network for filtering processing.

[0259] In this embodiment, by determining or obtaining the input of the neural network and / or the lookup table based on the eighth sub-input branch, the reconstruction transformation features of all channels of the reconstructed block of the chrominance component can be integrated together, and then filtering processing is performed through the neural network and / or the lookup table, thereby ensuring the effectiveness of the filtering processing.

[0260] Mode 9, a ninth sub-input branch for processing the prediction transformation feature on the chrominance component;

[0261] Optionally, the second input branch includes a ninth sub-input branch, and the predicted transformation features of the prediction block can be processed in the ninth sub-input branch. For example, channel splicing can be performed on all predicted transformation features on the chrominance component of at least one prediction block. Other processing operations can also be performed, such as converting them into a format that can be recognized by the neural network and / or lookup table, etc., to obtain corresponding input features so that they can be input into the lookup table and / or neural network for filtering processing.

[0262] In this embodiment, by determining or obtaining the input of the neural network and / or the lookup table based on the ninth sub-input branch, it is possible to integrate the predicted transformation features of all channels of the prediction block of the chrominance component, and then perform filtering processing through the neural network and / or the lookup table, thereby ensuring the effectiveness of the filtering processing.

[0263] Mode 10, a tenth sub-input branch for processing the derived transform feature on the chrominance component;

[0264] Optionally, the second input branch includes a tenth sub-input branch, and the derivative transformation features corresponding to the derivative blocks of the prediction block and / or the reconstructed block can be processed in the tenth sub-input branch. For example, channel splicing can be performed on all derivative transformation features on the chrominance component of at least one derivative block, and other processing operations can also be performed, such as converting them into a format that can be recognized by the neural network and / or lookup table, etc., to obtain the corresponding input features so that they can be input into the lookup table and / or neural network for filtering processing.

[0265] In this embodiment, by determining or obtaining the input of the neural network and / or the lookup table based on the tenth sub-input branch, it is possible to integrate the derivative transformation features of all channels of the derivative block of the chrominance component, and then perform filtering processing through the neural network and / or the lookup table, thereby ensuring the effectiveness of the filtering processing.

[0266] Mode 11, for processing the fourth input branch of the U component;

[0267] Optionally, the second input branch includes a fourth input branch.

[0268] Optionally, in the fourth input branch, channel splicing processing can be performed on the U component of the image block, such as performing channel splicing on the first feature of the derivative block on the U component and the second feature of the reconstructed block and / or the predicted block on the U component, and then determining or obtaining the input of the neural network and / or lookup table based on the result of the channel splicing.

[0269] In this embodiment, by determining or obtaining the input of the neural network and / or lookup table based on the fourth input branch for processing the U component, all information of the U component can be integrated together and then filtered through the neural network and / or lookup table, thereby ensuring the effectiveness of the filtering processing.

[0270] Mode 12 is used to process the fifth input branch of the V component.

[0271] Optionally, the second input branch includes a fifth input branch.

[0272] Optionally, in the fourth input branch, channel splicing processing can be performed on the V component of the image block, such as performing channel splicing on the first feature of the derivative block on the V component and the second feature of the reconstructed block and / or the predicted block on the V component, and then determining or obtaining the input of the neural network and / or lookup table based on the result of the channel splicing.

[0273] In this embodiment, by determining or obtaining the input of the neural network and / or lookup table based on the fourth input branch for processing the V component, all information of the V component can be integrated together and then filtered through the neural network and / or lookup table, thereby ensuring the effectiveness of the filtering processing.

[0274] Mode three is a third input branch for mixed processing of the luminance component and the chrominance component.

[0275] Optionally, in the third input branch, channel splicing processing can be performed on the luminance component and chrominance component of the image block, such as performing channel splicing on the first feature of the derivative block on the luminance component and the chrominance component, and the second feature of the reconstructed block and / or the predicted block on the luminance component and the chrominance component, and then determining or obtaining the input of the neural network and / or the lookup table based on the result of the channel splicing.

[0276] In this embodiment, by determining or obtaining the input of the neural network and / or the lookup table based on the third input branch for mixed processing of the luminance component and the chrominance component, all information of the luminance component and the chrominance component can be integrated together, and then filtered through the neural network and / or the lookup table, thereby ensuring the effectiveness of the filtering process.

[0277] Optionally, in step S11, channel stitching is performed based on multiple input branches and at least one image block feature, including at least one of steps b1 to b15:

[0278] Step b1: performing channel splicing on the predicted transformation feature of at least one luminance component, the reconstructed transformation feature of at least one luminance component, and the derived transformation feature of at least one luminance component according to the first input branch;

[0279] Optionally, various features of the brightness component may be processed in the first input branch.

[0280] Optionally, in the first input branch, predicted transformation features of at least one prediction block on the luminance component, reconstructed transformation features of at least one reconstructed block on the luminance component, derived transformation features of a derivative block of at least one reconstructed block on the luminance component, and derived transformation features of a derivative block of at least one prediction block on the luminance component can be determined.

[0281] Channel splicing can be performed on at least one predicted transformation feature, at least one reconstructed transformation feature, at least one derived transformation feature corresponding to a derivative block of a reconstructed block, and at least one derived transformation feature corresponding to a derivative block of a predicted block to determine or obtain at least one input feature, and then the at least one input feature is filtered through a neural network and / or a lookup table.

[0282] For example, Figure 9 As shown, taking the luminance component as an example, in the first input branch, the prediction block (PredY), the derivative block of the prediction block (PredY Overlap), the reconstruction block (RecY), and the derivative block of the reconstruction block (RecY Overlap) of the size of 144x144x1 are all subjected to 2x2DCT transformation (i.e., DCT-Ⅱ in the figure) and reconstructed into a 72x72x4 tensor, and then channel splicing is performed through the Rec module to obtain a 72x72x16 tensor (including multi-channel features), which is input into the neural network for model training, such as the convolution layer Conv3x3,16 performs group convolution processing to perform feature fusion through group convolution processing, and then outputs the filtered image block.

[0283] In this embodiment, by performing channel splicing on the predicted transformation features of at least one luminance component, the reconstructed transformation features of at least one luminance component, and the derived transformation features of at least one luminance component based on the first input branch to determine or obtain at least one input feature, and then processing the channel splicing result through a neural network and / or a lookup table, it is possible to effectively capture the transition features between image blocks on the luminance component, improve the filtering effect of the filtering process, and thereby improve the effect of video encoding and / or decoding.

[0284] Step b2: performing channel concatenation on the predicted transformation feature of at least one chroma component, the reconstructed transformation feature of at least one chroma component, and the derived transformation feature of at least one luminance component according to the second input branch;

[0285] Optionally, each feature of the chrominance component may be processed in the second input branch.

[0286] Optionally, in the second input branch, the predicted transformation characteristics of at least one prediction block on the chrominance component, the reconstructed transformation characteristics of at least one reconstructed block on the chrominance component, the derived transformation characteristics of the derivative block of at least one reconstructed block on the chrominance component, and the derived transformation characteristics of the derivative block of at least one prediction block on the chrominance component can be determined.

[0287] Optionally, channel splicing can be performed on at least one predicted transformation feature, at least one reconstructed transformation feature, a derivative transformation feature corresponding to a derivative block of at least one reconstructed block, and a derivative transformation feature corresponding to a derivative block of at least one predicted block to determine or obtain at least one input feature, and then the at least one input feature is filtered through a neural network and / or a lookup table.

[0288] In this embodiment, by performing channel splicing on the predicted transformation features of at least one chroma component, the reconstructed transformation features of at least one chroma component, and the derived transformation features of at least one luminance component based on the second input branch to determine or obtain at least one input feature, and then processing the channel splicing result through a neural network and / or a lookup table, it is possible to effectively capture the transition features between image blocks on the chroma component, improve the filtering effect of the filtering processing, and thereby improve the effect of video encoding and / or decoding.

[0289] Step b3: performing channel splicing on at least one of the predicted transformation feature of at least one luminance component and / or chrominance component, the derived transformation feature of at least one luminance component and / or chrominance component, and the reconstructed transformation feature of at least one luminance component and / or chrominance component according to the third input branch;

[0290] Optionally, the chrominance components and / or individual features of the chrominance components may be processed in the third input branch.

[0291] Optionally, in the second input branch, the predicted transformation features of at least one prediction block on the chrominance component, the predicted transformation features of at least one prediction block on the luminance component, the reconstructed transformation features of at least one reconstructed block on the chrominance component, the reconstructed transformation features of at least one reconstructed block on the luminance component, the derived transformation features of the derivative block of at least one reconstructed block on the chrominance component, the derived transformation features of the derivative block of at least one reconstructed block on the luminance component, the derived transformation features of the derivative block of at least one prediction block on the chrominance component, and the derived transformation features of the derivative block of at least one prediction block on the luminance component can be determined.

[0292] Optionally, channel splicing can be performed on at least one of the predicted transformation features of at least one luminance component and / or chrominance component, the derived transformation features of at least one luminance component and / or chrominance component (including the derived transformation features corresponding to the reconstructed block and / or the derived transformation features corresponding to the predicted block), and the reconstructed transformation features of at least one luminance component and / or chrominance component to determine or obtain at least one input feature, and then the at least one input feature is filtered through a neural network and / or a lookup table.

[0293] In this embodiment, by performing channel splicing on at least one of the predicted transformation features of at least one luminance component and / or chrominance component, the derived transformation features of at least one luminance component and / or chrominance component, and the reconstructed transformation features of at least one luminance component and / or chrominance component based on the third input branch to determine or obtain at least one input feature, and then processing the channel splicing result through a neural network and / or a lookup table, it is possible to effectively capture the transition features between image blocks on the chrominance component and the luminance component, improve the filtering effect of the filtering process, and thereby improve the effect of video encoding and / or decoding.

[0294] Step b4, performing channel splicing on the reconstructed transformation feature of at least one luminance component and the derived transformation feature of at least one luminance component according to the first sub-input branch;

[0295] Optionally, in the first sub-input branch, various features of the reconstructed block and the corresponding derivative block on the luminance component may be processed.

[0296] Optionally, a reconstruction transformation feature of the reconstructed block in the first sub-input branch in the luminance component and a derivative transformation feature of the derivative block of the reconstructed block in the luminance component may be determined.

[0297] Optionally, channel splicing can be performed on at least one reconstructed transformation feature and at least one derived transformation feature in the first sub-input branch to determine or obtain at least one input feature, and then the at least one input feature can be filtered through a neural network and / or a lookup table.

[0298] In this embodiment, by performing channel splicing on the reconstructed transformation features of at least one luminance component and the derived transformation features of at least one luminance component based on the first sub-input branch to determine or obtain at least one input feature, and then processing the channel splicing result through a neural network and / or a lookup table, it is possible to effectively capture the transition features between reconstructed blocks on the luminance component, improve the filtering effect of the filtering processing, and thereby improve the effect of video encoding and / or decoding.

[0299] Step b5: performing channel splicing on the predicted transformation feature of at least one luminance component and the derived transformation feature of at least one luminance component according to the second sub-input branch;

[0300] Optionally, in the second sub-input branch, various features of the prediction block and the corresponding derivative block on the luminance component may be processed.

[0301] Optionally, the predicted transformation feature of the prediction block in the second sub-input branch in the luminance component and the derived transformation feature of the derivative block of the prediction block in the luminance component may be determined.

[0302] Optionally, channel splicing can be performed on at least one predicted transformation feature and at least one derived transformation feature in the second sub-input branch to determine or obtain at least one input feature, and then the at least one input feature can be filtered through a neural network and / or a lookup table.

[0303] In this embodiment, by performing channel splicing on the predicted transformation features of at least one luminance component and the derived transformation features of at least one luminance component based on the second sub-input branch to determine or obtain at least one input feature, and then processing the channel splicing result through a neural network and / or a lookup table, it is possible to effectively capture the transition features between prediction blocks on the luminance component, improve the filtering effect of the filtering process, and thereby improve the effect of video encoding and / or decoding.

[0304] Step b6, performing channel splicing on the reconstructed transformation features of at least one brightness component according to the third sub-input branch;

[0305] Optionally, in the third sub-input branch, various features of the reconstructed block on the luminance component may be processed.

[0306] Optionally, the reconstruction transformation features of the reconstructed block in the third sub-input branch in all channels of the luminance component may be determined.

[0307] Optionally, the reconstructed transformation features of at least two channels can be channel-spliced in the third sub-input branch to determine or obtain at least one input feature, and then the at least one input feature can be filtered through a neural network and / or a lookup table.

[0308] In this embodiment, by performing channel splicing on the reconstructed transformation features of at least one luminance component based on the third sub-input branch to determine or obtain at least one input feature, and then processing the channel splicing result through a neural network and / or a lookup table, it is possible to achieve multi-channel unified processing of the reconstructed transformation features of the reconstructed block on the luminance component, thereby improving the filtering effect of the filtering process.

[0309] Step b7, performing channel splicing on the predicted transformation feature of at least one luminance component according to the fourth sub-input branch;

[0310] Optionally, in the fourth sub-input branch, various features of the prediction block on the luminance component may be processed.

[0311] Optionally, the prediction transformation features of the prediction block in the fourth sub-input branch in all channels of the luminance component may be determined.

[0312] Optionally, channel splicing can be performed on the predicted transformation features of at least two channels in the fourth sub-input branch to determine or obtain at least one input feature, and then the at least one input feature can be filtered through a neural network and / or a lookup table.

[0313] In this embodiment, by performing channel splicing on the predicted transformation features of at least one luminance component based on the fourth sub-input branch to determine or obtain at least one input feature, and then processing the channel splicing result through a neural network and / or a lookup table, it is possible to achieve multi-channel unified processing of the predicted transformation features of the prediction block on the luminance component, thereby improving the filtering effect of the filtering process.

[0314] Step b8, performing channel splicing on the derived transformation features of at least one luminance component according to the fifth sub-input branch;

[0315] Optionally, in the fifth sub-input branch, various features of the derivative block of the prediction block and / or the derivative block of the reconstructed block on the luminance component may be processed.

[0316] Optionally, the derivative transformation features of the derivative block in the fifth sub-input branch in all channels of the luminance component may be determined.

[0317] Optionally, the derived transformation features of at least two channels can be channel-spliced in the fifth sub-input branch to determine or obtain at least one input feature, and then the at least one input feature can be filtered through a neural network and / or a lookup table.

[0318] In this embodiment, by performing channel splicing on the derivative transformation features of at least one luminance component based on the fifth sub-input branch, and then processing the channel splicing results through a neural network and / or a lookup table, it is possible to perform multi-channel unified processing on the derivative transformation features of the derivative block on the luminance component, thereby improving the filtering effect of the filtering process.

[0319] Step b9: performing channel splicing on the reconstructed transformation feature of at least one chroma component and the derived transformation feature of at least one chroma component according to the sixth sub-input branch;

[0320] Optionally, in the sixth sub-input branch, various features of the reconstructed block and the corresponding derivative block on the chrominance component may be processed.

[0321] Optionally, a reconstruction transformation feature of the reconstructed block in the sixth sub-input branch in the chrominance component and a derivative transformation feature of the derivative block of the reconstructed block in the chrominance component may be determined.

[0322] Optionally, channel splicing can be performed on at least one reconstructed transformation feature and at least one derived transformation feature in the sixth sub-input branch to determine or obtain at least one input feature, and then the at least one input feature can be filtered through a neural network and / or a lookup table.

[0323] In this embodiment, by performing channel splicing on the reconstructed transformation features of at least one chrominance component and the derived transformation features of at least one chrominance component based on the sixth sub-input branch, and then processing the channel splicing results through a neural network and / or a lookup table, it is possible to effectively capture the transition features between reconstructed blocks on the chrominance components, improve the filtering effect of the filtering processing, and thereby improve the effect of video encoding and / or decoding.

[0324] Step b10: performing channel splicing on the predicted transformation feature of at least one chroma component and the derived transformation feature of at least one chroma component according to the seventh sub-input branch;

[0325] Optionally, in the seventh sub-input branch, various features of the prediction block and the corresponding derivative block on the chrominance component may be processed.

[0326] Optionally, the predicted transformation feature of the prediction block in the seventh sub-input branch in the chrominance component and the derived transformation feature of the derivative block of the prediction block in the chrominance component may be determined.

[0327] Optionally, channel splicing can be performed on at least one predicted transformation feature and at least one derived transformation feature in the seventh sub-input branch to determine or obtain at least one input feature, and then the at least one input feature can be filtered through a neural network and / or a lookup table.

[0328] In this embodiment, by performing channel splicing on the predicted transformation features of at least one chroma component and the derived transformation features of at least one chroma component based on the seventh sub-input branch to determine or obtain at least one input feature, and then processing the channel splicing result through a neural network and / or a lookup table, it is possible to effectively capture the transition features between prediction blocks on the chroma component, improve the filtering effect of the filtering process, and thereby improve the effect of video encoding and / or decoding.

[0329] Step b11, performing channel splicing on the reconstructed transformation feature of at least one chroma component according to the eighth sub-input branch;

[0330] Optionally, in the eighth sub-input branch, various features of the reconstructed block on the chrominance component may be processed.

[0331] Optionally, the reconstruction transformation features of the reconstructed block in the eighth sub-input branch in all channels of the chrominance components may be determined.

[0332] Optionally, the reconstructed transformation features of at least two channels can be channel-spliced in the eighth sub-input branch to determine or obtain at least one input feature, and then the at least one input feature can be filtered through a neural network and / or a lookup table.

[0333] In this embodiment, channel splicing is performed on the reconstructed transformation features of at least one chrominance component based on the eighth sub-input branch to determine or obtain at least one input feature, and the channel splicing result is then processed through a neural network and / or a lookup table, thereby achieving multi-channel unified processing of the reconstructed transformation features of the reconstructed block on the chrominance component, thereby improving the filtering effect of the filtering process.

[0334] Step b12, performing channel splicing on the predicted transformation feature of at least one chroma component according to the ninth sub-input branch;

[0335] Optionally, in the ninth sub-input branch, various features of the prediction block on the chrominance component may be processed.

[0336] Optionally, the prediction transformation features of the prediction block in the ninth sub-input branch in all channels of the chrominance components may be determined.

[0337] Optionally, channel splicing can be performed on the predicted transformation features of at least two channels in the ninth sub-input branch to determine or obtain at least one input feature, and then the at least one input feature can be filtered through a neural network and / or a lookup table.

[0338] In this embodiment, by performing channel splicing on the predicted transformation features of at least one chrominance component based on the ninth sub-input branch to determine or obtain at least one input feature, and then processing the channel splicing result through a neural network and / or a lookup table, it is possible to achieve multi-channel unified processing of the predicted transformation features of the prediction block on the chrominance component, thereby improving the filtering effect of the filtering process.

[0339] Step b13, performing channel splicing on the derived transformation feature of at least one chroma component according to the tenth sub-input branch;

[0340] Optionally, in the tenth sub-input branch, various features of the derivative block of the prediction block and / or the derivative block of the reconstructed block on the chrominance component may be processed.

[0341] Optionally, derivative transformation features of the derivative block in the tenth sub-input branch on all channels of the chrominance components may be determined.

[0342] Optionally, the derived transformation features of at least two channels can be channel-spliced in the tenth sub-input branch to determine or obtain at least one input feature, and then the at least one input feature can be filtered through a neural network and / or a lookup table.

[0343] In this embodiment, by performing channel splicing on the derivative transformation features of at least one chrominance component based on the tenth sub-input branch, and then processing the channel splicing results through a neural network and / or a lookup table, it is possible to perform multi-channel unified processing on the derivative transformation features of the derivative block on the luminance component, thereby improving the filtering effect of the filtering process.

[0344] Step b14: performing channel splicing on the predicted transformation feature of at least one U component, the reconstructed transformation feature of at least one U component, and the derived transformation feature of at least one U component according to the fourth input branch;

[0345] Optionally, various features of the U component may be processed in the fourth input branch.

[0346] Optionally, in the fourth input branch, predicted transformation features of at least one prediction block on the U component, reconstructed transformation features of at least one reconstructed block on the U component, derived transformation features of a derivative block of at least one reconstructed block on the U component, and derived transformation features of a derivative block of at least one prediction block on the U component can be determined.

[0347] Optionally, channel splicing can be performed on at least one predicted transformation feature, at least one reconstructed transformation feature, a derivative transformation feature corresponding to a derivative block of at least one reconstructed block, and a derivative transformation feature corresponding to a derivative block of at least one predicted block to determine or obtain at least one input feature, and then the at least one input feature is filtered through a neural network and / or a lookup table.

[0348] In this embodiment, by performing channel splicing on the predicted transformation features of at least one U component, the reconstructed transformation features of at least one U component, and the derived transformation features of at least one U component based on the fourth input branch to determine or obtain at least one input feature, and then processing the channel splicing result through a neural network and / or a lookup table, it is possible to effectively capture the transition features between image blocks on the U component, improve the filtering effect of the filtering processing, and thereby improve the effect of video encoding and / or decoding.

[0349] Step b15: performing channel concatenation on the predicted transformation feature of at least one V component, the reconstructed transformation feature of at least one V component, and the derived transformation feature of at least one V component according to the fifth input branch.

[0350] Optionally, various features of the V component may be processed in the fourth input branch.

[0351] Optionally, in the fourth input branch, the predicted transformation characteristics of at least one prediction block on the V component, the reconstructed transformation characteristics of at least one reconstructed block on the V component, the derived transformation characteristics of the derivative block of at least one reconstructed block on the V component, and the derived transformation characteristics of the derivative block of at least one prediction block on the V component can be determined.

[0352] Optionally, channel splicing can be performed on at least one predicted transformation feature, at least one reconstructed transformation feature, a derivative transformation feature corresponding to a derivative block of at least one reconstructed block, and a derivative transformation feature corresponding to a derivative block of at least one predicted block to determine or obtain at least one input feature, and then the at least one input feature is filtered through a neural network and / or a lookup table.

[0353] In this embodiment, by performing channel splicing on the predicted transformation features of at least one V component, the reconstructed transformation features of at least one V component, and the derived transformation features of at least one V component based on the fifth input branch to determine or obtain at least one input feature, and then processing the channel splicing results through a neural network and / or a lookup table, it is possible to effectively capture the transition features between image blocks on the V component, improve the filtering effect of the filtering processing, and thereby improve the effect of video encoding and / or decoding.

[0354] Step S12: filtering the at least one input feature according to the neural network and / or the lookup table.

[0355] Optionally, at least one index may be determined or obtained based on at least one input feature, and the at least one index may be input into a lookup table for search, and the filtered image block may be determined or obtained based on the search result. Correspondences between various features and indices may be pre-set, and the index corresponding to the at least one input feature may be determined based on the correspondences.

[0356] Optionally, at least one input feature may be input into a neural network for filtering.

[0357] In this embodiment, by filtering at least one input feature based on a neural network and / or a lookup table, the advantages of the neural network and / or the lookup table can be combined, and the transition characteristics between image blocks can be effectively captured through derivative blocks, reconstructed blocks and / or prediction blocks, which can improve the filtering effect of the filtering process and thereby improve the effect of video encoding and / or decoding.

[0358] Third embodiment

[0359] Based on any of the above embodiments, a third embodiment is proposed.

[0360] In this embodiment, the neural network includes a group convolution module for dividing at least one input feature into at least one group for convolution processing, and / or at least one group includes at least one of the methods four to nineteen.

[0361] Optionally, since the neural network includes a grouped convolution module, after at least one input feature is input into the neural network, the at least one input feature can be grouped and processed using the grouped convolution module in the neural network, and / or the grouping and processing can be performed according to a one-to-one correspondence with the input branches, or according to the number of channels, without any limitation here.

[0362] Optionally, a convolution module may be used to perform convolution processing within each group, or a partial convolution module may be used to perform partial convolution processing, which is not limited here.

[0363] Mode 4, for processing at least one group of features corresponding to the first input branch in at least one input feature;

[0364] Optionally, in the neural network, at least one group corresponding to the first input branch can be determined, and within the at least one group corresponding to the first input branch, the reconstructed transformation features of the reconstructed block on the luminance component, the derived transformation features of the derivative block of the reconstructed block on the luminance component, the predicted transformation features of the predicted block on the luminance component, and the derived transformation features of the derivative block of the predicted block on the luminance component can be convolved through a convolution module.

[0365] In this embodiment, by performing group convolution processing in a neural network and performing convolution processing using at least one group of features corresponding to the first input branch in processing at least one input feature, the effective progress of filtering is ensured, and the amount of calculation is effectively reduced while maintaining the feature expression capability.

[0366] Mode 5, for processing at least one group of features corresponding to the second input branch in at least one input feature;

[0367] Optionally, in the neural network, at least one group corresponding to the second input branch can be determined, and within the at least one group corresponding to the second input branch, the reconstructed transformation features of the reconstructed block on the chrominance component, the derived transformation features of the derivative block of the reconstructed block on the chrominance component, the predicted transformation features of the predicted block on the chrominance component, and the derived transformation features of the derivative block of the predicted block on the chrominance component can be convolved through a convolution module.

[0368] In this embodiment, by performing group convolution processing in a neural network and performing convolution processing using at least one group of features corresponding to the second input branch in processing at least one input feature, the effective progress of filtering is ensured, and the amount of calculation is effectively reduced while maintaining the feature expression capability.

[0369] Mode six, for processing at least one group of features corresponding to the third input branch in at least one input feature;

[0370] Optionally, in the neural network, at least one group corresponding to the third input branch can be determined, and within the at least one group corresponding to the third input branch, the reconstructed transformation features of the reconstructed block on the chrominance component, the derived transformation features of the derivative block of the reconstructed block on the chrominance component, the predicted transformation features of the prediction block on the chrominance component, the derived transformation features of the derivative block of the prediction block on the chrominance component, the reconstructed transformation features of the reconstructed block on the luminance component, the derived transformation features of the derivative block of the prediction block on the luminance component, the predicted transformation features of the derivative block of the prediction block on the luminance component, and the derived transformation features of the derivative block of the prediction block on the luminance component are convolved through a convolution module.

[0371] In this embodiment, by performing group convolution processing in a neural network and performing convolution processing using at least one group of features corresponding to the third input branch in processing at least one input feature, the effective progress of filtering is ensured, and the amount of calculation is effectively reduced while maintaining the feature expression capability.

[0372] Mode seven, for processing at least one group of features corresponding to the first sub-input branch in at least one input feature;

[0373] Optionally, in the neural network, at least one group corresponding to the first sub-input branch may be determined, and within the at least one group corresponding to the first sub-input branch, a convolution module may be used to perform convolution processing on the reconstructed transformation features of the reconstructed block on the luminance component and the derived transformation features of the derived block of the reconstructed block on the luminance component. Convolution processing may also be performed using a partial convolution module, i.e., convolution processing may be performed on only a portion of the reconstructed transformation features and / or the derived transformation features.

[0374] In this embodiment, by performing group convolution processing in a neural network and performing convolution processing using at least one group of features corresponding to the first sub-input branch in at least one input feature, the effective progress of filtering is ensured, and the amount of calculation is effectively reduced while maintaining the feature expression capability.

[0375] Mode eight, for processing at least one group of features corresponding to the second sub-input branch in at least one input feature;

[0376] Optionally, in the neural network, at least one group corresponding to the second sub-input branch may be determined, and within the at least one group corresponding to the second sub-input branch, a convolution module may be used to perform convolution processing on the predicted transform features on the luminance component of the prediction block and the derived transform features on the luminance component of a derivative block of the prediction block. Convolution processing may also be performed using a partial convolution module, i.e., convolution processing may be performed on only a portion of the predicted transform features and / or the derived transform features.

[0377] In this embodiment, by performing group convolution processing in a neural network and performing convolution processing using at least one group of features corresponding to the second sub-input branch in at least one input feature, the effective progress of filtering is ensured, and the amount of calculation is effectively reduced while maintaining the feature expression capability.

[0378] Mode nine, for processing at least one group of features corresponding to the third sub-input branch in at least one input feature;

[0379] Optionally, in the neural network, at least one group corresponding to the third sub-input branch may be determined, and within the at least one group corresponding to the third sub-input branch, a convolution module may be used to perform convolution processing on the reconstructed transform features on the luminance component of the reconstructed block. Furthermore, a partial convolution module may be used to perform convolution processing, i.e., convolution processing may be performed on only a portion of the predicted transform features and / or the derived transform features.

[0380] In this embodiment, by performing group convolution processing in a neural network and performing convolution processing using at least one group of features corresponding to the third sub-input branch in at least one input feature, the effective progress of filtering is ensured, and the amount of calculation is effectively reduced while maintaining the feature expression capability.

[0381] Method 10, for processing at least one group of features corresponding to the fourth sub-input branch in at least one input feature;

[0382] Optionally, in the neural network, at least one group corresponding to the fourth sub-input branch may be determined, and within the at least one group corresponding to the fourth sub-input branch, a convolution module may be used to perform convolution processing on the predicted transformation features of the prediction block on the luma component. Convolution processing may also be performed using a partial convolution module, i.e., convolution processing may be performed on only a portion of the predicted transformation features and / or the derived transformation features.

[0383] In this embodiment, by performing group convolution processing in a neural network and performing convolution processing using at least one group of features corresponding to the fourth sub-input branch in at least one input feature, the effective progress of filtering is ensured, and the amount of calculation is effectively reduced while maintaining the feature expression capability.

[0384] Mode 11, for processing at least one group of features corresponding to the fifth sub-input branch in at least one input feature;

[0385] Optionally, in the neural network, at least one group corresponding to the fifth sub-input branch may be determined, and within the at least one group corresponding to the fifth sub-input branch, a convolution module may be used to perform convolution processing on the derived transform features of the derivative block of the prediction block on the luminance component and / or the derived transform features of the derivative block of the reconstructed block on the luminance component. Alternatively, a partial convolution module may be used to perform convolution processing, i.e., convolution processing may be performed on only a portion of the prediction transform features and / or the derived transform features.

[0386] In this embodiment, by performing group convolution processing in a neural network and performing convolution processing using at least one group of features corresponding to the fifth sub-input branch in at least one input feature, the effective progress of filtering is ensured, and the amount of calculation is effectively reduced while maintaining the feature expression capability.

[0387] Mode 12, for processing at least one group of features corresponding to the sixth sub-input branch in at least one input feature;

[0388] Optionally, in the neural network, at least one group corresponding to the sixth sub-input branch may be determined, and within the at least one group corresponding to the sixth sub-input branch, a convolution module may be used to perform convolution processing on the reconstructed transform features of the reconstructed block on the chrominance components and the derived transform features of the derived block of the reconstructed block on the chrominance components. Convolution processing may also be performed using a partial convolution module, i.e., convolution processing may be performed on only a portion of the reconstructed transform features and / or the derived transform features.

[0389] In this embodiment, by performing group convolution processing in a neural network and performing convolution processing using at least one group of features corresponding to the sixth sub-input branch in at least one input feature, the effective progress of filtering is ensured, and the amount of calculation is effectively reduced while maintaining the feature expression capability.

[0390] Mode 13, for processing at least one group of features corresponding to the seventh sub-input branch in at least one input feature;

[0391] Optionally, in the neural network, at least one group corresponding to the seventh sub-input branch may be determined, and within the at least one group corresponding to the seventh sub-input branch, a convolution module may be used to perform convolution processing on the predicted transform features of the prediction block on the chrominance components and the derived transform features of the derivative block of the prediction block on the chrominance components. Convolution processing may also be performed using a partial convolution module, i.e., convolution processing may be performed on only a portion of the predicted transform features and / or the derived transform features.

[0392] In this embodiment, by performing group convolution processing in a neural network and performing convolution processing using at least one group of features corresponding to the seventh sub-input branch in at least one input feature, the effective progress of filtering is ensured, and the amount of calculation is effectively reduced while maintaining the feature expression capability.

[0393] Mode 14, for processing at least one group of features corresponding to the eighth sub-input branch in at least one input feature;

[0394] Optionally, in the neural network, at least one group corresponding to the eighth sub-input branch may be determined, and within the at least one group corresponding to the eighth sub-input branch, a convolution module may be used to perform convolution processing on the reconstructed transform features of the reconstructed block on the chrominance components. Furthermore, a partial convolution module may be used to perform convolution processing, i.e., convolution processing may be performed on only a portion of the predicted transform features and / or the derived transform features.

[0395] In this embodiment, by performing group convolution processing in a neural network and performing convolution processing using at least one group of features corresponding to the eighth sub-input branch in at least one input feature, the effective progress of filtering is ensured, and the amount of calculation is effectively reduced while maintaining the feature expression capability.

[0396] Mode 15, for processing at least one group of features corresponding to the ninth sub-input branch in at least one input feature;

[0397] Optionally, in the neural network, at least one group corresponding to the ninth sub-input branch may be determined, and within the at least one group corresponding to the ninth sub-input branch, a convolution module may be used to perform convolution processing on the predicted transformation features of the prediction block on the chrominance components. Convolution processing may also be performed using a partial convolution module, i.e., convolution processing may be performed on only a portion of the predicted transformation features and / or the derived transformation features.

[0398] In this embodiment, by performing group convolution processing in a neural network and performing convolution processing using at least one group of features corresponding to the ninth sub-input branch in at least one input feature, the effective progress of filtering is ensured, and the amount of calculation is effectively reduced while maintaining the feature expression capability.

[0399] Mode 16, for processing at least one group of features corresponding to the tenth sub-input branch in at least one input feature;

[0400] Optionally, in the neural network, at least one group corresponding to the tenth sub-input branch may be determined, and within the at least one group corresponding to the tenth sub-input branch, a convolution module may be used to perform convolution processing on the derived transform features of the derivative block of the prediction block on the chrominance components and / or the derived transform features of the derivative block of the reconstructed block on the chrominance components. Alternatively, a partial convolution module may be used to perform convolution processing, i.e., convolution processing may be performed on only a portion of the prediction transform features and / or the derived transform features.

[0401] In this embodiment, by performing group convolution processing in a neural network and performing convolution processing using at least one group of features corresponding to the tenth sub-input branch in at least one input feature, the effective progress of filtering is ensured, and the amount of calculation is effectively reduced while maintaining the feature expression capability.

[0402] Mode 17, for processing at least one group of features corresponding to the fourth input branch in at least one input feature;

[0403] Optionally, in the neural network, at least one group corresponding to the fourth input branch may be determined, and within the at least one group corresponding to the fourth input branch, a convolution module may be used to perform convolution processing on the reconstructed transform features of the reconstructed block on the U component, the derived transform features of the derived block of the reconstructed block on the U component, the predicted transform features of the predicted block on the U component, and the derived transform features of the derived block of the predicted block on the U component. Convolution processing may also be performed using a partial convolution module, i.e., convolution processing may be performed on only a portion of the predicted transform features and / or the derived transform features.

[0404] In this embodiment, by performing group convolution processing in a neural network and performing convolution processing using at least one group of features corresponding to the fourth input branch in processing at least one input feature, the effective progress of filtering is ensured, and the amount of calculation is effectively reduced while maintaining the feature expression capability.

[0405] Mode 18, for processing at least one group of features corresponding to the fifth input branch in at least one input feature;

[0406] Optionally, in the neural network, at least one group corresponding to the fifth input branch may be determined, and within the at least one group corresponding to the fifth input branch, a convolution module may be used to perform convolution processing on the reconstructed transform features of the reconstructed block on the V component, the derived transform features of the derived block of the reconstructed block on the V component, the predicted transform features of the predicted block on the V component, and the derived transform features of the derived block of the predicted block on the V component. Convolution processing may also be performed using a partial convolution module, i.e., convolution processing may be performed on only a portion of the predicted transform features and / or the derived transform features.

[0407] In this embodiment, by performing group convolution processing in a neural network and performing convolution processing using at least one group of features corresponding to the fifth input branch in processing at least one input feature, the effective progress of filtering is ensured, and the amount of calculation is effectively reduced while maintaining the feature expression capability.

[0408] Mode 19, including at least one group of at least one channel among the multiple channels of each input branch.

[0409] Optionally, the input branch is at least one item in multiple input branches corresponding to at least one input feature.

[0410] Optionally, grouping can be performed in a cross-channel manner, for example, in each input branch of the multi-input branch, image features of at least one channel are selected as a separate group for convolution processing.

[0411] Optionally, one channel-dimensional input feature (such as predicted transformation feature, reconstructed transformation feature, derived transformation feature, etc.) can be selected from each input branch in the multiple input branches and put into a group for convolution processing until the input features of each channel dimension in each input branch complete the corresponding convolution processing in their respective corresponding groups.

[0412] For example, if the multi-input branch includes a first input branch and a second input branch, the first input branch has input feature 1 with a channel dimension of 1, input feature 2 with a channel dimension of 2, and input feature 3 with a channel dimension of 3, and the second input branch has input feature 4 with a channel dimension of 4, input feature 5 with a channel dimension of 5, and input feature 6 with a channel dimension of 6. Then, in the neural network, the grouped convolution module can be used to perform grouped convolution processing on each input feature, such as the first group processing input feature 1 with a channel dimension of 1 and input feature 4 with a channel dimension of 4, the second group processing input feature 2 with a channel dimension of 2 and input feature 5 with a channel dimension of 5, and the third group processing input feature 3 with a channel dimension of 3 and input feature 6 with a channel dimension of 6.

[0413] In this embodiment, by performing group convolution processing in a neural network and performing convolution processing using at least one group of at least one channel among multiple channels of each input branch, the effective implementation of the filtering process is guaranteed, and the amount of calculation is effectively reduced while maintaining the feature expression capability.

[0414] Fourth embodiment

[0415] Based on any of the above embodiments, a fourth embodiment is proposed.

[0416] In this embodiment, the image processing method further includes at least one of the following methods 20 to 29:

[0417] Mode 20, the grouped convolution module includes at least one convolution module corresponding to a group, and the convolution modules corresponding to at least one group are independent of each other;

[0418] Optionally, in the neural network, the grouped convolution module includes a convolution module corresponding to at least one group, that is, the convolution module may exist in some groups, and may not exist in some groups. The outputs of all groups are then channel-spliced and subsequently processed until the neural network outputs the corresponding target image block after filtering.

[0419] Optionally, the convolution modules corresponding to at least one group are independent of each other. There may be multiple groups, each applying its own corresponding convolution module to perform corresponding convolution processing, and / or the convolution modules in multiple groups can perform convolution processing in parallel.

[0420] In this embodiment, the group convolution module of the neural network includes at least one convolution module corresponding to a group, and the convolution modules corresponding to at least one group are independent of each other, which can ensure the effective implementation of the group convolution.

[0421] Method 21: at least two groups contain the same number of input channels;

[0422] Optionally, in the neural network, the number of input channels corresponding to the input features input to some or all groups is equal, that is, some or all groups process different input features with the same number of channels.

[0423] In this embodiment, at least two groups in the neural network include an equal number of input channels, which can achieve effective group convolution.

[0424] Method 22: The convolutional layer parameters of at least two groups are the same;

[0425] Optionally, in the neural network, the convolution layer parameters of the convolution modules in some groups may be the same.

[0426] Optionally, the convolution layer parameters of the convolution modules in some groups may be different.

[0427] In this embodiment, the convolution layer parameters of at least two groups in the neural network are the same, so that the effective implementation of group convolution can be achieved.

[0428] Mode 23, the grouped convolution module sequentially includes a precursor layer for each group, a first channel shuffle layer, and a subsequent convolution module for each group;

[0429] Mode 24: The input of the first channel shuffle layer is the output of at least two groups of predecessor layers, and the output of the first channel shuffle layer is the input of at least two subsequent groups of convolution modules;

[0430] In mode 25, the precursor layer includes at least one convolution module for performing convolution processing on at least one group of features;

[0431] Mode 26: The first channel shuffling layer is used to perform cross-group permutation of the output feature map of the predecessor layer in the channel dimension by using the first channel rearrangement rule;

[0432] Optionally, in the neural network, multiple convolution modules and a first channel shuffle layer can be set in each group. The first channel shuffle layer can be set in the middle of multiple convolution modules, and the convolution module located before the first channel shuffle layer in the same group can be used as a precursor layer, and the convolution module located after the first channel shuffle layer can be used as a subsequent convolution module.

[0433] Optionally, the convolution module may include at least one convolution layer, such as a 1x1 convolution layer.

[0434] Optionally, the precursor layer may include a 1x1 convolution layer, and a smaller lookup table may be used to replace at least one convolution module included in the precursor layer. For example, the index corresponding to at least one group of features is input into the lookup table for search, and then cross-group permutation in the channel dimension is performed through the first channel shuffling layer based on the search result.

[0435] Optionally, within at least one group, at least one convolution module in the precursor layer performs deep feature fusion and transformation on all features within at least one group to obtain higher quality features, which are then processed through the first channel shuffling layer.

[0436] Optionally, the precursor layer can deeply fuse the intra-group features, deeply fuse and refine all channel information in the same group, generate high-quality intra-group feature representations, and perform dimensionality reduction processing on the generated high-quality features to reduce the computational complexity of the first channel shuffling layer.

[0437] Optionally, first-pass shuffling layers are interspersed in each group.

[0438] Optionally, the first channel shuffling layer can promote the flow and fusion of information between different channels by rearranging the channel order of the input feature map. The channel dimension of the input feature map can be reshaped into two dimensions, such as the number of convolution groups and the number of channels contained in each convolution group. The two reshaped dimensions are then transposed, and the number of convolution groups (such as the number of multiple groups) and the number of channels contained in each convolution group (such as each group) are swapped. The transposed channel dimension is then flattened and restored to the original number of channels, thereby improving the expressiveness of the features without increasing the computational cost.

[0439] Optionally, for any group, at least one input feature can be input into the predecessor layer for convolution processing, and then the output of the predecessor layer can be input into the first channel shuffling layer for cross-group permutation operation in the channel dimension, and then the output of the first channel shuffling layer can be input into the subsequent convolution module for convolution processing, until the neural network outputs the target image block after filtering.

[0440] Optionally, in the first channel shuffling layer, the output results of the predecessor layer of each group can be permuted across groups in the channel dimension according to the first channel rearrangement rule. For example, the features corresponding to the first channel of the first group and the features corresponding to the first channel of the second group can be permuted across groups.

[0441] Optionally, the output feature map of the predecessor layer includes the output result of the predecessor layer.

[0442] Optionally, the first channel rearrangement rule may be a pre-set channel dimension rearrangement rule, which is not limited here.

[0443] In this embodiment, each group in the neural network is sequentially provided with a precursor layer including at least one convolution module, a first channel shuffling layer, and a subsequent convolution module of each group, and / or when performing group convolution processing, the precursor layer, the first channel shuffling layer, and the subsequent convolution module are sequentially processed, and / or the first channel shuffling layer can use the first channel rearrangement rule to permute the output feature map of the predecessor layer across groups in the channel dimension, which can improve the filtering effect of filtering processing on at least one image block.

[0444] Mode 27: The grouped convolution module includes a second channel shuffle layer and a convolution module for each group, and the input of the convolution module for each group is the output of the second channel shuffle layer;

[0445] Mode 28: The second channel shuffling layer is used to perform cross-group permutation on the channel dimension of an input feature map including features of at least two groups by using a second channel rearrangement rule;

[0446] Optionally, in the neural network, at least one convolutional module and a second channel shuffling layer may be provided within each group. The second channel shuffling layer may be provided before the at least one convolutional module, and the second channel shuffling layer may be interspersed within each group. Optionally, the second channel shuffling layer may refer to the first channel shuffling layer described above, and will not be repeated here.

[0447] Optionally, for any group, at least one input feature of that group can be input to a second channel shuffling layer. Within the second channel shuffling layer, the input feature map comprising features of at least one group is permuted across groups along the channel dimension using a second channel permutation rule. The output of the second channel shuffling layer is then input to a subsequent convolution module of each group for convolution processing until the neural network outputs a filtered target image block.

[0448] Optionally, each group corresponds to an input feature map, and the input feature map includes input features input to at least one group.

[0449] Optionally, the second channel reordering rule may be the same as or different from the first channel reordering rule, which is not limited here.

[0450] Optionally, the second channel rearrangement rule may be a channel dimension rearrangement rule set in advance.

[0451] In this embodiment, the neural network includes a second channel shuffling layer and a convolution module for each group in sequence, and the second channel shuffling layer can perform cross-group permutation on the channel dimension of the input feature map including features of at least one group through a second channel rearrangement rule, thereby improving the filtering effect of filtering at least one image block.

[0452] Method 29, the grouped convolution module includes, in sequence, a second channel shuffling layer, a precursor layer of each group, a first channel shuffling layer, and a subsequent convolution module of each group.

[0453] Optionally, in the neural network, a second channel shuffle layer, a precursor layer, a first channel shuffle layer, and a subsequent convolution module may be sequentially arranged in each group.

[0454] Optionally, for any group, at least one input feature input to the group can be input to the second channel shuffle layer, and within the second channel shuffle layer, the input feature map (including at least one input feature) including at least one group of features is subjected to the above-mentioned cross-group permutation in the channel dimension by the second channel rearrangement rule. The output of the second channel shuffle layer is then input to the predecessor layer of each group for convolution processing, and the output of the predecessor layer (i.e., the output feature map) is then input to the first channel shuffle layer, and the output feature map of the predecessor layer is subjected to cross-group permutation processing in the channel dimension by the first channel rearrangement rule, and the output of the first channel shuffle layer is then input to the subsequent convolution module for convolution processing, until the neural network outputs the target image block after filtering.

[0455] In this embodiment, the neural network includes, in sequence, a second channel shuffling layer, a predecessor layer of each group, a first channel shuffling layer, and a subsequent convolution module of each group, and / or can perform cross-group permutation operations multiple times in the channel dimension through the first channel shuffling layer and the second channel shuffling layer to improve the filtering effect of filtering processing on at least one image block.

[0456] Fifth embodiment

[0457] Based on any of the above embodiments, a fifth embodiment is proposed.

[0458] In this embodiment, step S12 includes at least one of steps c1 to c5:

[0459] Step c1, determining or obtaining at least one first intermediate value based on the neural network and at least one input feature, and performing filtering processing on at least one image block based on the at least one first intermediate value and at least one lookup table;

[0460] Optionally, the first intermediate value may be data generated by an intermediate process during the filtering process, may be a filtered pixel generated after at least one filtering process, or may be another value. Optionally, the intermediate process may be a process of performing pre-filtering processing on the pixel to be filtered (e.g., filtering processing using a neural network or a lookup table).

[0461] Optionally, at least one input feature can be input into the neural network for filtering processing, and the output is a first intermediate value. For example, at least one of the quantization parameter, boundary strength, characterization interval, position information of the pixel to be filtered, the pixel value to be filtered, the size information of the image block, the filtering information of the neighboring block, the filtering information of the non-neighboring block, the filtering information of the cross-component block, the filtering information of the co-located block, the filtering information of the time domain block, the filtering information of the default block and the filtering information of the candidate block can be input into the neural network for filtering processing, and the output is a first intermediate value.

[0462] Optionally, the first intermediate value may be directly used as a filtering pixel, and the target image block may be determined or generated based on the filtering pixel.

[0463] Optionally, the first intermediate value may be processed to determine or generate a target image block, for example, by inputting the first intermediate value into at least one lookup table for lookup to determine a filtered pixel, and determining or generating the target image block based on the at least one filtered pixel. Optionally, the lookup table may include a correspondence between the first intermediate value and the filtered pixel.

[0464] Optionally, the at least one lookup table may be a multi-layer lookup table. The first intermediate value may be converted into an index and input into the multi-layer lookup table for parallel or serial lookup, and a filtered pixel may be output, and a target image block may be determined or generated based on the at least one filtered pixel.

[0465] In this embodiment, by determining or generating at least one first intermediate value based on a neural network and at least one input feature, determining or generating a target image block based on at least one first intermediate value and at least one lookup table, and performing filtering processing together using the neural network and the lookup table, the filtering effect can be improved, thereby improving the effect of video encoding and / or decoding.

[0466] Step c2, determining or obtaining at least one second intermediate value based on the at least one input feature and the at least one lookup table, and performing filtering processing on the at least one image block based on the neural network and the at least one second intermediate value;

[0467] Optionally, the second intermediate value may be data generated during an intermediate process of the filtering process, may be a filtered pixel generated after at least one filtering process, or may be another value. Optionally, the second intermediate value may be the same as or different from the first intermediate value.

[0468] Optionally, at least one selected lookup table may be determined from a plurality of lookup tables based on the at least one input feature, and the at least one input feature may be input into the at least one lookup table for search and determination to obtain the at least one second intermediate value. Optionally, the lookup table includes a correspondence between the at least one input feature and the second intermediate value, and / or an index may be assigned to each correspondence in the lookup table.

[0469] Optionally, at least one lookup table may be a multi-layer lookup table. The pixel values to be filtered of at least one image block may be converted into an index and input into the multi-layer lookup table for parallel or serial search, and at least one second intermediate value may be output. Optionally, based on at least one of a quantization parameter, a boundary strength, a characterization parameter, position information of the pixel to be filtered, image block size information, filtering information of neighboring blocks, filtering information of non-neighboring blocks, filtering information of cross-component blocks, filtering information of co-located blocks, filtering information of time-domain blocks, filtering information of default blocks, and filtering information of candidate blocks, the index converted from the pixel values to be filtered of at least one image block may be updated (e.g., the index value may be increased or decreased), and the updated index may be input into the multi-layer lookup table for parallel or serial search, and at least one second intermediate value may be output.

[0470] Optionally, a lookup table can be screened based on at least one of the quantization parameter, boundary strength, characterization interval, position information of the pixel to be filtered, the pixel value to be filtered, size information of the image block, filtering information of the neighboring block, filtering information of the non-neighboring block, filtering information of the cross-component block, filtering information of the same-position block, filtering information of the time domain block, filtering information of the default block and filtering information of the candidate block, and the pixel value to be filtered can be converted into an index and input into the screened lookup table for search to determine the filtered pixel after filtering, and output it as the second intermediate value.

[0471] Optionally, at least one second intermediate value may be input into a neural network, and filtered pixels may be outputted, and a target image block may be determined or generated based on the filtered pixels.

[0472] Optionally, the neural network may be determined by screening based on at least one input feature.

[0473] Optionally, the image block including at least one second intermediate value may be input into a neural network for filtering, and a target image block including filtered pixels may be output.

[0474] In this embodiment, at least one second intermediate value is obtained by searching in at least one lookup table based on at least one input feature, a target image block is determined or generated based on a neural network and the at least one second intermediate value, and filtering is performed together using the neural network and the lookup table. This can improve the filtering effect and thereby improve the effect of video encoding and / or decoding.

[0475] Step c3, determining or generating at least one index based on at least one input feature, and performing filtering processing on at least one image block based on the at least one index and at least one lookup table;

[0476] Optionally, at least one index may be determined by at least one input feature in at least one image block, and a method for determining the index may refer to at least one of Methods 31 to 37 in the following embodiments.

[0477] Optionally, the index can be a number, an array, an identifier, a label, etc.

[0478] Optionally, it is possible to determine whether it is necessary to update the index of the pixel value to be filtered of at least one image block based on at least one of the quantization parameter, boundary strength, characterization parameter, position information of the pixel to be filtered, size information of the image block, filtering information of the neighboring block, filtering information of the non-neighboring block, filtering information of the cross-component block, filtering information of the same-position block, filtering information of the time domain block, filtering information of the default block and filtering information of the candidate block. If necessary, the converted index can be increased or decreased to obtain at least one updated index.

[0479] Optionally, after determining at least one index, the at least one index may be input into at least one lookup table for search to determine and output a corresponding filtered pixel, and the target image block may be determined or generated based on the output of the at least one filtered pixel.

[0480] Alternatively, after determining or generating at least one index, such as a first index, based on at least one input feature, the first index may be input into a first-level lookup table in a multi-layer lookup table for search, obtaining a first search result. A second index may then be determined or obtained based on the first search result, and the second index may be input into a second-level lookup table in the multi-layer lookup table for search, until a final lookup table outputs a filtered pixel or a target image block containing the filtered pixel. Optionally, the at least one lookup table may include a multi-layer lookup table.

[0481] In this embodiment, at least one index is determined or generated based on at least one input feature, and at least one lookup table is used to determine or generate a target image block. By using the lookup table for filtering, the complexity of the filtering process can be reduced, thereby improving the efficiency of video encoding and / or decoding.

[0482] Step c4, determining or obtaining at least one third intermediate value based on the neural network and the at least one input feature, determining or obtaining at least one fourth intermediate value based on the at least one third intermediate value and at least one lookup table, and performing filtering processing on at least one image block based on the neural network and the at least one fourth intermediate value;

[0483] Optionally, the third intermediate value and / or the fourth intermediate value may be data generated in an intermediate processing process during the filtering process, may be filtered pixels generated after at least one filtering process, or may be other values.

[0484] Optionally, the fourth intermediate value, and / or the third intermediate value, and / or the second intermediate value, and / or the first intermediate value may be the same or different.

[0485] Alternatively, a neural network may be determined or selected based on at least one input feature (e.g., reconstructed transformation features, predicted transformation features, and / or derived transformation features of different components), and the at least one input feature may be input into the neural network for filtering, with the third intermediate value being output. Alternatively, at least one image block and the input features of the at least one image block may be input into the neural network for filtering, with the image block being output after a single filtering process. Pixels in the image block after the single filtering process may be used as the third intermediate value.

[0486] Optionally, the at least one third intermediate value can be input into at least one lookup table for search and determination to generate at least one fourth intermediate value. Optionally, the at least one third intermediate value can be converted into an index (e.g., directly using the third intermediate value as the index, or transforming the third intermediate value to obtain the index), and the index is then input into the at least one lookup table for search, with the at least one fourth intermediate value being output. Optionally, the lookup table includes a correspondence between the third intermediate values and the fourth intermediate values, and / or indexes can be set in the lookup table, with each correspondence corresponding to an index.

[0487] Optionally, the at least one lookup table may be a multi-layer lookup table, and the at least one third intermediate value may be input into the multi-layer lookup table for serial or parallel lookup, and the at least one fourth intermediate value may be output.

[0488] Optionally, at least one fourth intermediate value may be input into a neural network for model training, and a target image block including at least one filtered pixel may be output.

[0489] In this embodiment, at least one input feature is filtered based on a neural network, a lookup table, and a neural network architecture to determine or generate a target image block. By using the neural network and the lookup table together for filtering, the filtering effect can be improved, thereby improving the effect of video encoding and / or decoding.

[0490] Step c5: determining or obtaining at least one fifth intermediate value based on at least one input feature and at least one lookup table, determining or obtaining at least one sixth intermediate value based on at least one fifth intermediate value and a neural network, and filtering at least one image block based on at least one sixth intermediate value and at least one lookup table.

[0491] Optionally, the fifth intermediate value and / or the sixth intermediate value may be data generated during an intermediate process of the filtering process, may be filtered pixels generated after at least one filtering process, or may be other values. Optionally, the sixth intermediate value, and / or the fifth intermediate value, and / or the fourth intermediate value, and / or the third intermediate value, and / or the second intermediate value, and / or the first intermediate value may be the same or different.

[0492] Alternatively, at least one selected lookup table may be determined from a plurality of lookup tables based on at least one input feature (e.g., reconstructed transformation features, and / or predicted transformation features, and / or derived transformation features of different components), and the at least one input feature may be input into the at least one lookup table to determine the at least one fifth intermediate value. Alternatively, an index may be determined based on the at least one input feature, and the index may be input into the at least one lookup table to determine the at least one fifth intermediate value.

[0493] Optionally, the lookup table includes a correspondence between at least one input feature and the fifth intermediate value, and / or an index may be set for each correspondence in the lookup table.

[0494] Optionally, at least one lookup table may be a multi-layer lookup table. The first input feature may be converted into an index and input into the multi-layer lookup table for parallel or serial lookup, and at least one fifth intermediate value may be output.

[0495] Optionally, the at least one fifth intermediate value may be input into a neural network for filtering, and at least one sixth intermediate value may be output. Alternatively, an image block including the at least one fifth intermediate value may be input into a neural network for filtering, and at least one sixth intermediate value may be output.

[0496] Alternatively, the at least one sixth intermediate value may be input into at least one lookup table for search and determination of a filter pixel, and the target image block may be determined or generated based on the filter pixel. Alternatively, an index may be determined based on the at least one sixth intermediate value and input into at least one lookup table for search and determination of the filter pixel.

[0497] Optionally, the lookup table may further include a correspondence between at least one sixth intermediate value and a filtered pixel, and / or an index may be set for each correspondence in the lookup table.

[0498] Optionally, at least one sixth intermediate value may be input into a multi-layer lookup table for parallel or serial lookup, and a filtered pixel may be output, or a target image block containing the filtered pixel may be output.

[0499] In this embodiment, at least one input feature is filtered based on the architecture of a lookup table, a neural network, and a lookup table to determine or generate a target image block. By using the neural network and the lookup table together for filtering, the filtering effect can be improved, thereby improving the effect of video encoding and / or decoding.

[0500] Sixth embodiment

[0501] The present application also provides a processing device, referring to Figure 10 , the processing device includes:

[0502] The processing module A10 is configured to perform filtering processing on at least one image block based on multiple input branches, a neural network, and / or a lookup table.

[0503] Optionally, the image block feature comprises at least one of a reconstruction transformation feature of a reconstructed block of at least one image block, a prediction transformation feature of a prediction block of at least one image block, and a derivative transformation feature of a derivative block of at least one image block;

[0504] Optionally, the processing module A10 is configured to:

[0505] Perform channel splicing based on multiple input branches and at least one image block feature to determine or obtain input features;

[0506] Filtering is performed on at least one input feature according to a neural network and / or a lookup table.

[0507] Optionally, the multi-input branch includes at least one of the following:

[0508] a first input branch for processing a luminance component;

[0509] a second input branch for processing chrominance components;

[0510] The third input branch is used for mixing the luminance component and the chrominance component.

[0511] Optionally, the first input branch includes at least one of the following: a first sub-input branch for processing the reconstructed transformation features and the derived transformation features on the luminance component, a second sub-input branch for processing the predicted transformation features and the derived transformation features on the luminance component, a third sub-input branch for processing the reconstructed transformation features on the luminance component, a fourth sub-input branch for processing the predicted transformation features on the luminance component, and a fifth sub-input branch for processing the derived transformation features on the luminance component; and / or,

[0512] The second input branch includes at least one of the following: a sixth sub-input branch for processing the reconstructed transformation features and the derived transformation features on the chroma component, a seventh sub-input branch for processing the predicted transformation features and the derived transformation features on the chroma component, an eighth sub-input branch for processing the reconstructed transformation features on the chroma component, a ninth sub-input branch for processing the predicted transformation features on the chroma component, a tenth sub-input branch for processing the derived transformation features on the chroma component, a fourth input branch for processing the U component, and a fifth input branch for processing the V component.

[0513] Optionally, the processing module A10 is further configured to perform at least one of the following:

[0514] performing channel concatenation on the predicted transformation feature of at least one luminance component, the reconstructed transformation feature of at least one luminance component, and the derived transformation feature of at least one luminance component according to the first input branch;

[0515] performing channel concatenation on the predicted transform feature of at least one chroma component, the reconstructed transform feature of at least one chroma component, and the derived transform feature of at least one luminance component according to the second input branch;

[0516] performing channel splicing on at least one of a predicted transform feature of at least one luminance component and / or chrominance component, a derived transform feature of at least one luminance component and / or chrominance component, and a reconstructed transform feature of at least one luminance component and / or chrominance component according to the third input branch;

[0517] performing channel splicing on the reconstructed transformation feature of at least one luminance component and the derived transformation feature of at least one luminance component according to the first sub-input branch;

[0518] performing channel concatenation on the predicted transformation feature of at least one luminance component and the derived transformation feature of at least one luminance component according to the second sub-input branch;

[0519] performing channel splicing on the reconstructed transformation features of at least one brightness component according to the third sub-input branch;

[0520] performing channel splicing on the predicted transformation feature of at least one luminance component according to the fourth sub-input branch;

[0521] performing channel splicing on the derived transformation features of at least one luminance component according to the fifth sub-input branch;

[0522] performing channel splicing on the reconstructed transformation feature of at least one chroma component and the derived transformation feature of at least one chroma component according to the sixth sub-input branch;

[0523] performing channel concatenation on the predicted transformation feature of at least one chroma component and the derived transformation feature of at least one chroma component according to the seventh sub-input branch;

[0524] performing channel splicing on the reconstructed transformation feature of at least one chroma component according to the eighth sub-input branch;

[0525] performing channel splicing on the predicted transformation feature of at least one chrominance component according to the ninth sub-input branch;

[0526] performing channel splicing on the derived transform feature of at least one chroma component according to the tenth sub-input branch;

[0527] performing channel concatenation on the predicted transformation feature of at least one U component, the reconstructed transformation feature of at least one U component, and the derived transformation feature of at least one U component according to the fourth input branch;

[0528] According to the fifth input branch, channel concatenation is performed on the predicted transformation feature of the at least one V component, the reconstructed transformation feature of the at least one V component, and the derived transformation feature of the at least one V component.

[0529] Optionally, the neural network includes a group convolution module for dividing at least one input feature into at least one group for convolution processing;

[0530] Optionally, at least one group includes at least one of the following:

[0531] for processing at least one group of features corresponding to the first input branch in at least one input feature;

[0532] for processing at least one group of features corresponding to the second input branch in at least one input feature;

[0533] for processing at least one group of features corresponding to a third input branch in at least one input feature;

[0534] for processing at least one group of features corresponding to the first sub-input branch in at least one input feature;

[0535] for processing at least one group of features corresponding to the second sub-input branch in at least one input feature;

[0536] for processing at least one group of features corresponding to the third sub-input branch in at least one input feature;

[0537] for processing at least one group of features corresponding to a fourth sub-input branch in at least one input feature;

[0538] for processing at least one group of features corresponding to the fifth sub-input branch in at least one input feature;

[0539] for processing at least one group of features corresponding to a sixth sub-input branch in at least one input feature;

[0540] for processing at least one group of features corresponding to a seventh sub-input branch in at least one input feature;

[0541] for processing at least one group of features corresponding to an eighth sub-input branch in at least one input feature;

[0542] for processing at least one group of features corresponding to a ninth sub-input branch in at least one input feature;

[0543] for processing at least one group of features corresponding to a tenth sub-input branch in at least one input feature;

[0544] for processing at least one group of features corresponding to a fourth input branch in at least one input feature;

[0545] for processing at least one group of features corresponding to a fifth input branch in at least one input feature;

[0546] At least one group includes at least one channel of the plurality of channels of each input branch.

[0547] Optionally, the input branch is at least one item in multiple input branches corresponding to at least one input feature.

[0548] Optionally, the processing module A10 is further configured to perform at least one of the following:

[0549] The grouped convolution module includes at least one convolution module corresponding to a group, and the convolution modules corresponding to at least one group are independent of each other;

[0550] At least two groups contain the same number of input channels;

[0551] The convolutional layer parameters of at least two groups are the same;

[0552] The grouped convolution module sequentially includes the precursor layer of each group, the first channel shuffle layer, and the subsequent convolution module of each group;

[0553] The input of the first channel shuffle layer is the output of at least two groups of predecessor layers, and the output of the first channel shuffle layer is the input of at least two subsequent groups of convolution modules;

[0554] The precursor layer includes at least one convolution module for performing convolution processing on at least one set of features;

[0555] The first channel shuffling layer is used to permute the output feature maps of the predecessor layer across groups in the channel dimension through the first channel rearrangement rule;

[0556] The grouped convolution module includes a second channel shuffle layer and a convolution module for each group, and the input of the convolution module for each group is the output of the second channel shuffle layer;

[0557] The second channel shuffling layer is used to perform cross-group permutation on the channel dimension of the input feature map including features of at least two groups through a second channel rearrangement rule;

[0558] The grouped convolution module sequentially includes the second channel shuffle layer, the predecessor layer of each group, the first channel shuffle layer, and the subsequent convolution module of each group.

[0559] Optionally, the processing module A10 is further configured to perform at least one of the following:

[0560] Determine or obtain at least one first intermediate value based on the neural network and at least one input feature, and perform filtering processing on at least one image block based on the at least one first intermediate value and at least one lookup table;

[0561] Determine or obtain at least one second intermediate value based on at least one input feature and at least one lookup table, and perform filtering on at least one image block based on the neural network and the at least one second intermediate value;

[0562] Determine or generate at least one index according to at least one input feature, and perform filtering processing on at least one image block according to the at least one index and at least one lookup table;

[0563] Determining or obtaining at least one third intermediate value based on the neural network and the at least one input feature, determining or obtaining at least one fourth intermediate value based on the at least one third intermediate value and at least one lookup table, and filtering the at least one image block based on the neural network and the at least one fourth intermediate value;

[0564] At least one fifth intermediate value is determined or obtained based on at least one input feature and at least one lookup table, at least one sixth intermediate value is determined or obtained based on the at least one fifth intermediate value and a neural network, and filtering is performed on at least one image block based on the at least one sixth intermediate value and the at least one lookup table.

[0565] Optionally, the derivative block is determined or obtained by at least one of the following:

[0566] a cropping result of cropping a reconstructed block and / or a predicted block of at least one image block;

[0567] a filling result of filling a reconstructed block and / or a predicted block of at least one image block;

[0568] an update result of performing pixel update on a reconstructed block and / or a predicted block of at least one image block;

[0569] a translation result of performing pixel translation on a reconstructed block and / or a predicted block of at least one image block;

[0570] An editing result of performing pixel editing on a reconstructed block and / or a prediction block of at least one image block.

[0571] The processing device provided in the embodiment of the present application has similar implementation principles and beneficial effects to the technical solutions shown in the above-mentioned corresponding method embodiments, and will not be described in detail here.

[0572] An embodiment of the present application further provides a processing device, including a memory and a processor. The memory stores an image processing program, and when the image processing program is executed by the processor, the steps of the image processing method in any of the above embodiments are implemented.

[0573] An embodiment of the present application further provides a storage medium on which an image processing program is stored. When the image processing program is executed by a processor, the steps of the image processing method in any of the above embodiments are implemented.

[0574] In the embodiments of the processing device and storage medium provided in this application, all technical features of any of the above-mentioned image processing method embodiments may be included. The expanded and explained contents of the specification are basically the same as those of the embodiments of the above-mentioned methods and will not be repeated here.

[0575] An embodiment of the present application further provides a computer program product, which includes computer program code. When the computer program code runs on a computer, the computer executes the methods in the various possible implementation modes described above.

[0576] An embodiment of the present application also provides a chip, including a memory and a processor, wherein the memory is used to store computer programs, and the processor is used to call and run the computer programs from the memory, so that a device equipped with the chip executes the methods in the various possible implementation modes as described above.

[0577] It is understood that the above scenarios are merely examples and do not limit the application scenarios of the technical solutions provided in the embodiments of this application. The technical solutions of this application can also be applied to other scenarios. For example, those skilled in the art will appreciate that with the evolution of system architecture and the emergence of new business scenarios, the technical solutions provided in the embodiments of this application will also be applicable to similar technical problems.

[0578] The serial numbers of the above embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.

[0579] The steps in the method of the embodiment of the present application can be adjusted in order, combined and deleted according to actual needs.

[0580] The units in the device of the embodiment of the present application can be merged, divided and deleted according to actual needs.

[0581] In this application, for the same or similar terminology, technical solutions and / or application scenario descriptions, they are generally only described in detail the first time they appear. When they appear again later, for the sake of brevity, they are generally not repeated. When understanding the technical solutions and other contents of this application, for the same or similar terminology, technical solutions and / or application scenario descriptions that are not described in detail later, reference can be made to the relevant detailed descriptions before. In this application, the descriptions of each embodiment have their own emphasis. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments. The various technical features of the technical solution of this application can be combined arbitrarily. In order to make the description concise, not all possible combinations of the various technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.

[0582] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus the necessary general hardware platform, and of course can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) as mentioned above, and includes a number of instructions for enabling a terminal device (which can be a mobile phone, computer, server, controlled terminal, or network device, etc.) to execute the method of each embodiment of the present application.

[0583] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When software is used for implementation, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function according to the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a storage medium or transmitted from one storage medium to another storage medium. For example, the computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) means. The storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more available media integrations. The available medium can be a magnetic medium (e.g., a floppy disk, a storage disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state storage disk Solid State Disk (SSD)).

[0584] The above are only preferred embodiments of the present application and do not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made using the contents of the present application specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present application.

Claims

1. An image processing method, characterized in that: Including steps: S1, performing filtering processing on at least one image block according to multiple input branches, a neural network and / or a lookup table.

2. The image processing method according to claim 1, wherein: The image block feature includes at least one of a reconstructed transform feature of a reconstructed block of at least one image block, a predicted transform feature of a predicted block of at least one image block, and a derived transform feature of a derived block of at least one image block; and / or, step S1 includes the steps of: S11, performing channel splicing based on multiple input branches and at least one image block feature to determine or obtain input features; S12, performing filtering processing on at least one input feature according to a neural network and / or a lookup table.

3. The image processing method according to claim 2, wherein: A multi-input branch includes at least one of the following: a first input branch for processing a luminance component; a second input branch for processing chrominance components; The third input branch is used for mixing the luminance component and the chrominance component.

4. The image processing method according to claim 3, wherein: The first input branch includes at least one of the following: a first sub-input branch for processing the reconstructed transformation feature and the derived transformation feature on the luma component, a second sub-input branch for processing the predicted transformation feature and the derived transformation feature on the luma component, a third sub-input branch for processing the reconstructed transformation feature on the luma component, a fourth sub-input branch for processing the predicted transformation feature on the luma component, and a fifth sub-input branch for processing the derived transformation feature on the luma component; and / or, The second input branch includes at least one of the following: a sixth sub-input branch for processing the reconstructed transformation features and the derived transformation features on the chroma component, a seventh sub-input branch for processing the predicted transformation features and the derived transformation features on the chroma component, an eighth sub-input branch for processing the reconstructed transformation features on the chroma component, a ninth sub-input branch for processing the predicted transformation features on the chroma component, a tenth sub-input branch for processing the derived transformation features on the chroma component, a fourth input branch for processing the U component, and a fifth input branch for processing the V component.

5. The image processing method according to claim 4, wherein: Channel stitching is performed based on multiple input branches and at least one image block feature, including at least one of the following: performing channel concatenation on the predicted transformation feature of at least one luminance component, the reconstructed transformation feature of at least one luminance component, and the derived transformation feature of at least one luminance component according to the first input branch; performing channel concatenation on the predicted transform feature of at least one chroma component, the reconstructed transform feature of at least one chroma component, and the derived transform feature of at least one luminance component according to the second input branch; performing channel splicing on at least one of a predicted transform feature of at least one luminance component and / or chrominance component, a derived transform feature of at least one luminance component and / or chrominance component, and a reconstructed transform feature of at least one luminance component and / or chrominance component according to the third input branch; performing channel splicing on the reconstructed transformation feature of at least one luminance component and the derived transformation feature of at least one luminance component according to the first sub-input branch; performing channel concatenation on the predicted transformation feature of at least one luminance component and the derived transformation feature of at least one luminance component according to the second sub-input branch; performing channel splicing on the reconstructed transformation features of at least one brightness component according to the third sub-input branch; performing channel splicing on the predicted transformation feature of at least one luminance component according to the fourth sub-input branch; performing channel splicing on the derived transformation features of at least one luminance component according to the fifth sub-input branch; performing channel splicing on the reconstructed transformation feature of at least one chroma component and the derived transformation feature of at least one chroma component according to the sixth sub-input branch; performing channel concatenation on the predicted transformation feature of at least one chroma component and the derived transformation feature of at least one chroma component according to the seventh sub-input branch; performing channel splicing on the reconstructed transformation feature of at least one chroma component according to the eighth sub-input branch; performing channel splicing on the predicted transformation feature of at least one chrominance component according to the ninth sub-input branch; performing channel splicing on the derived transform feature of at least one chroma component according to the tenth sub-input branch; performing channel concatenation on the predicted transformation feature of at least one U component, the reconstructed transformation feature of at least one U component, and the derived transformation feature of at least one U component according to the fourth input branch; According to the fifth input branch, channel concatenation is performed on the predicted transformation feature of the at least one V component, the reconstructed transformation feature of the at least one V component, and the derived transformation feature of the at least one V component.

6. The image processing method according to claim 4, wherein: The neural network includes a grouped convolution module for dividing at least one input feature into at least one group for convolution processing, and / or at least one group includes at least one of the following: for processing at least one group of features corresponding to the first input branch in at least one input feature; for processing at least one group of features corresponding to the second input branch in at least one input feature; for processing at least one group of features corresponding to a third input branch in at least one input feature; for processing at least one group of features corresponding to the first sub-input branch in at least one input feature; for processing at least one group of features corresponding to the second sub-input branch in at least one input feature; for processing at least one group of features corresponding to the third sub-input branch in at least one input feature; for processing at least one group of features corresponding to a fourth sub-input branch in at least one input feature; for processing at least one group of features corresponding to the fifth sub-input branch in at least one input feature; for processing at least one group of features corresponding to a sixth sub-input branch in at least one input feature; for processing at least one group of features corresponding to a seventh sub-input branch in at least one input feature; for processing at least one group of features corresponding to an eighth sub-input branch in at least one input feature; for processing at least one group of features corresponding to a ninth sub-input branch in at least one input feature; for processing at least one group of features corresponding to a tenth sub-input branch in at least one input feature; for processing at least one group of features corresponding to a fourth input branch in at least one input feature; for processing at least one group of features corresponding to a fifth input branch in at least one input feature; At least one group includes at least one channel of the plurality of channels of each input branch.

7. The image processing method according to claim 6, wherein: Also include at least one of the following: The grouped convolution module includes at least one convolution module corresponding to a group, and the convolution modules corresponding to at least one group are independent of each other; At least two groups contain the same number of input channels; The convolutional layer parameters of at least two groups are the same; The input branch is at least one item in the multiple input branches corresponding to at least one input feature; The grouped convolution module sequentially includes the precursor layer of each group, the first channel shuffle layer, and the subsequent convolution module of each group; The input of the first channel shuffle layer is the output of at least two groups of predecessor layers, and the output of the first channel shuffle layer is the input of at least two subsequent groups of convolution modules; The precursor layer includes at least one convolution module for performing convolution processing on at least one set of features; The first channel shuffling layer is used to permute the output feature maps of the predecessor layer across groups in the channel dimension through the first channel rearrangement rule; The grouped convolution module includes a second channel shuffle layer and a convolution module for each group, and the input of the convolution module for each group is the output of the second channel shuffle layer; The second channel shuffling layer is used to perform cross-group permutation on the channel dimension of the input feature map including features of at least two groups through a second channel rearrangement rule; The grouped convolution module sequentially includes the second channel shuffle layer, the predecessor layer of each group, the first channel shuffle layer, and the subsequent convolution module of each group.

8. The image processing method according to any one of claims 2 to 7, wherein: Step S12 includes at least one of the following: Determine or obtain at least one first intermediate value based on the neural network and at least one input feature, and perform filtering processing on at least one image block based on the at least one first intermediate value and at least one lookup table; Determine or obtain at least one second intermediate value based on at least one input feature and at least one lookup table, and perform filtering on at least one image block based on the neural network and the at least one second intermediate value; Determine or generate at least one index according to at least one input feature, and perform filtering processing on at least one image block according to the at least one index and at least one lookup table; determining or obtaining at least one third intermediate value based on the neural network and the at least one input feature, determining or obtaining at least one fourth intermediate value based on the at least one third intermediate value and at least one lookup table, and performing filtering on at least one image block based on the neural network and the at least one fourth intermediate value; At least one fifth intermediate value is determined or obtained based on at least one input feature and at least one lookup table, at least one sixth intermediate value is determined or obtained based on the at least one fifth intermediate value and a neural network, and filtering is performed on at least one image block based on the at least one sixth intermediate value and the at least one lookup table.

9. The image processing method according to any one of claims 2 to 7, wherein: A derived block is determined or obtained by at least one of the following: a cropping result of cropping a reconstructed block and / or a predicted block of at least one image block; a filling result of filling a reconstructed block and / or a predicted block of at least one image block; an update result of performing pixel update on a reconstructed block and / or a predicted block of at least one image block; a translation result of performing pixel translation on a reconstructed block and / or a predicted block of at least one image block; An editing result of performing pixel editing on a reconstructed block and / or a prediction block of at least one image block.

10. A processing device, characterized in that include: A memory and a processor, wherein an image processing program is stored in the memory, and when the image processing program is executed by the processor, the steps of the image processing method according to any one of claims 1 to 9 are implemented.

11. A storage medium, characterized in that: The storage medium stores a computer program, which, when executed by a processor, implements the steps of the image processing method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Image processing method, image processing device and equipment

    CN111311629A

  • Neural network structure searching method and device, image processing method and device

    CN112215332A

  • Image processing method and device, image training method and device, channel shuffling method and device

    CN112927174A

  • Image reconstruction method, image coding and decoding method and related equipment

    CN114943643A

  • Code stream decoding and coding based on neural network

    CN116648912A

Cited By

  • Image processing method, processing equipment and storage medium

    CN121056641A