Image processing method, processing equipment and storage medium
By using the spatial domain characteristics and spatial attention weight of the reference block to predict the image block, the problem that the neural network cannot focus on the key image block characteristics is solved, and a more efficient image block prediction effect is achieved.
Patent Information
- Application Number
- CN202510832169.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-19
- Publication Date
- 2025-09-02
- Estimated Expiration
- 2045-06-19
AI Technical Summary
In image processing, the neural network cannot focus on key image block features during prediction processing, resulting in poor prediction results.
By predicting the image blocks based on the spatial characteristics of the reference block, the key image block features are determined using the spatial characteristics and spatial attention weights, and prediction processing is performed.
It improves the accuracy and efficiency of image block prediction processing, focuses on key image block features, and improves the prediction effect.
Smart Images

Figure CN120583237A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image processing technology, and in particular to an image processing method, processing device and storage medium. Background Art
[0002] Existing high-efficiency video coding frameworks, such as Neural Network Based Video Coding (NNVC) and / or Enhanced Compression Model (ECM), propose a video frame encoding technology to improve coding performance without significantly increasing computational complexity.
[0003] During the process of conceiving and implementing this application, the inventors discovered at least the following problems:
[0004] During the prediction processing stage of the encoding and decoding process, neural networks can be used for prediction processing, such as NN Intra Prediction (neural network intra-frame prediction) and NN Inter Prediction (neural network inter-frame prediction). However, due to the large number of image block features, the neural network needs to process multiple image block features and cannot focus on key image block features, resulting in poor prediction results, which needs to be improved.
[0005] The preceding description is intended to provide general background information and does not necessarily constitute prior art. Summary of the Invention
[0006] In response to the above technical problems, the present application provides an image processing method, a processing device and a storage medium, aiming to solve the technical problem of how to improve the prediction effect of prediction processing.
[0007] The present application provides an image processing method, which can be applied to a processing device, comprising the steps of:
[0008] S1, performing prediction processing on at least one image block according to spatial domain features of at least one reference block.
[0009] Optionally, the spatial domain feature is determined or obtained according to at least one of the following:
[0010] Statistical feature values of a reference pixel and at least one neighboring pixel in at least one reference block;
[0011] Statistical feature values of at least two non-neighbor pixels in at least one reference block;
[0012] a statistical feature value of a reference pixel in at least one reference block and at least one of the above neighboring pixel, at least one left neighboring pixel, at least one right neighboring pixel, and at least one below neighboring pixel of the reference pixel;
[0013] a weighted statistical feature value of a reference pixel and pixels adjacent to the reference pixel in at least one reference block;
[0014] The statistical characteristic value of at least one pixel in the outermost layer of a window centered on the reference pixel, wherein the size parameter of the window is less than or equal to the size parameter of the reference block;
[0015] Statistical eigenvalues of eight pixels that are not adjacent to the reference pixel, and the eight pixels are not adjacent to each other;
[0016] a statistical characteristic value of a reference sub-block of at least one reference block including at least one of at least one reference pixel, at least one neighboring pixel, and at least one non-neighboring pixel;
[0017] The at least one reference block includes a statistical characteristic value of a reference area of at least one of at least one reference pixel, at least one neighboring pixel, and at least one non-neighboring pixel.
[0018] Optionally, the reference block is determined or obtained according to at least one of the following:
[0019] At least one of an upper adjacent pixel, an upper non-adjacent pixel, a left adjacent pixel, a left non-adjacent pixel, an upper left adjacent pixel, and an upper left non-adjacent pixel of the image block;
[0020] At least one of a neighbor block, a non-neighbor block, a cross-component block, a co-located block, a time domain block, and a default block corresponding to the image block;
[0021] at least one of the width, height, block size, and block area of the image block;
[0022] a candidate block determined or generated by a candidate motion vector or a candidate block vector of the image block;
[0023] If the first information of the image block satisfies the first condition, the reference block is the first reference block;
[0024] If the first information of the image block satisfies the second condition, the reference block is the second reference block.
[0025] Optionally, step S1 includes at least one of the following:
[0026] Determining or obtaining at least one spatial attention weight based on at least one spatial feature and at least one lookup table, and performing prediction processing on at least one image block based on the at least one spatial attention weight;
[0027] Determining or obtaining at least one spatial attention weight based on at least one spatial feature and at least one neural network, and performing prediction processing on at least one image block based on the at least one spatial attention weight;
[0028] At least one spatial attention weight is determined or obtained based on at least one spatial feature and at least one activation function, and prediction processing is performed on at least one image block based on the at least one spatial attention weight.
[0029] Optionally, the image processing method further includes at least one of the following:
[0030] The spatial attention weights corresponding to at least two first intermediate elements in the intermediate element block corresponding to at least one reference block are the same and / or the spatial attention weights corresponding to the reference pixels are different;
[0031] The spatial attention weight corresponding to at least one first intermediate element in the intermediate element block corresponding to at least one reference block is greater than the spatial attention weight corresponding to another reference pixel;
[0032] The spatial attention weight corresponding to the intermediate element region containing at least one first intermediate element in the intermediate element block corresponding to at least one reference block is greater than the spatial attention weight corresponding to the intermediate element region containing at least one other first intermediate element;
[0033] The spatial attention weight corresponding to the intermediate element sub-block containing at least one first intermediate element in the intermediate element block corresponding to at least one reference block is greater than the spatial attention weight corresponding to the intermediate element sub-block containing at least one other first intermediate element.
[0034] Optionally, the image processing method further includes at least one of the following:
[0035] Determining or obtaining at least one intermediate element block based on at least one neural network, a lookup table structure, at least one item in the lookup table, and at least one reference block;
[0036] The lookup table structure includes at least one of the following:
[0037] at least one lookup table;
[0038] at least one neural network module;
[0039] At least two search branches, at least one search branch having the same input as another search branch;
[0040] at least two search branches, the input of at least one search branch being determined or obtained based on the output of another search branch;
[0041] At least two search branches, the at least two search branches being located in at least one channel corresponding to at least one reference pixel or a first reference intermediate element;
[0042] At least two search branches, the search tables of the at least two search branches are searched in a parallel search and / or serial search manner;
[0043] at least two lookup tables, wherein the input of at least one lookup table is determined or obtained according to the type, size parameter and / or input range of the other lookup table;
[0044] At least two lookup tables, the at least two lookup tables being located in at least one channel corresponding to at least one reference pixel or a first reference intermediate element;
[0045] At least two lookup tables, wherein the lookup mode of the at least two lookup tables in the same channel is parallel search and / or serial search;
[0046] At least two lookup tables, wherein the lookup tables in at least two channels are searched in parallel and / or serially;
[0047] at least one search branch and at least one neural network module, wherein an input of the at least one neural network module is the same as an input of the at least one search branch;
[0048] at least one search branch and at least one neural network module, wherein an input of the at least one neural network module is determined or obtained based on an output of the at least one search branch;
[0049] At least one lookup table and at least one neural network module, wherein an input of the at least one neural network module is determined or obtained according to a type, size parameter and / or input range of the at least one lookup table;
[0050] at least two lookup tables, wherein the output ranges of the at least two lookup tables are the same and / or the output ranges of the lookup tables are different;
[0051] at least two lookup tables, at least two of which have the same input range and / or different input ranges;
[0052] At least one first residual module based on a lookup table, the first residual module comprising a first branch including at least one lookup table structure, an adder, and a second branch including a short-circuit structure;
[0053] At least one second residual module based on a lookup table, a first input element of the second residual module passes through a first branch containing at least one lookup table structure to obtain a fourth output element, and a first input element of the third residual module passes through a second branch containing a short-circuit structure to obtain a fifth output element that is the same as the first input element, and the fourth output element and the fifth output element are weightedly added by an adder to obtain a sixth output element.
[0054] Optionally, performing prediction processing on at least one image block according to at least one spatial attention weight includes at least one of the following:
[0055] Determine or obtain at least one predicted block or predicted intermediate element according to a result of performing spatial attention enhancement on a first intermediate element in at least one intermediate element block according to at least one spatial attention weight;
[0056] Determine or obtain at least one predicted block or predicted intermediate element based on a result of performing spatial attention enhancement on at least one intermediate element region in at least one intermediate element block according to at least one spatial attention weight;
[0057] At least one predicted block or predicted intermediate element is determined or obtained based on the result of performing spatial attention enhancement on at least one intermediate element sub-block in at least one intermediate element block.
[0058] Optionally, the image processing method further includes: determining or obtaining spatial domain features of at least one reference block after image preprocessing.
[0059] Optionally, the image preprocessing includes at least one of the following:
[0060] Image sharpening;
[0061] Gradient calculation;
[0062] Transformation;
[0063] Detection is performed based on the detection operator.
[0064] The present application also provides a processing device, comprising:
[0065] The processing module is used to perform prediction processing on at least one image block according to the spatial domain characteristics of at least one reference block.
[0066] The present application also provides a processing device, comprising: a memory and a processor, wherein an image processing program is stored in the memory, and when the image processing program is executed by the processor, the steps of any of the above-mentioned image processing methods are implemented.
[0067] The present application also provides a storage medium storing a computer program, wherein the computer program, when executed by a processor, implements the steps of any of the above-mentioned image processing methods.
[0068] As described above, the image processing method of the present application can be applied to a processing device, including: performing predictive processing on at least one image block based on the spatial characteristics of at least one reference block. Through the technical solution of the present application, when performing predictive processing on at least one image block, the spatial characteristics of at least one reference block are comprehensively considered to determine key image block features that need to be processed in the reference block. This allows the prediction processing to focus on key image block features, thereby improving the prediction effect of the prediction processing using at least one reference block. BRIEF DESCRIPTION OF THE DRAWINGS
[0069] The accompanying drawings herein are incorporated into and constitute a part of the specification, illustrate embodiments consistent with the present application, and together with the specification, are used to explain the principles of the present application. In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following is a brief introduction to the drawings required for describing the embodiments. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without inventive work.
[0070] Figure 1 A schematic diagram of the hardware structure of a mobile terminal for implementing various embodiments of the present application;
[0071] Figure 2 A communication network system architecture diagram provided in an embodiment of the present application;
[0072] Figure 3 A schematic diagram of the hardware structure of a controller 140 provided in this application;
[0073] Figure 4 A schematic diagram of the hardware structure of a network node 150 provided in this application;
[0074] Figure 5 is a flowchart of an image processing method according to the first embodiment;
[0075] Figure 6 It is a flowchart of encoding and decoding in the image processing method;
[0076] Figure 7 It is a schematic diagram of spatial feature extraction of an image block in an image processing method;
[0077] Figure 8 It is a schematic diagram of a processing module of a processing device.
[0078] The purpose of this application, its features, and advantages will be further described in conjunction with the embodiments and with reference to the accompanying drawings. The accompanying drawings illustrate specific embodiments of this application, which will be described in more detail below. These drawings and the accompanying text are not intended to limit the scope of the present application in any way, but rather to illustrate the concepts of this application to those skilled in the art by reference to specific embodiments. DETAILED DESCRIPTION
[0079] Exemplary embodiments will be described in detail herein, with examples illustrated in the accompanying drawings. In the following description, when referring to the drawings, identical numerals in different figures represent identical or similar elements, unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all embodiments consistent with the present application. Rather, they are merely examples of apparatus and methods consistent with certain aspects of the present application, as detailed in the appended claims.
[0080] It should be noted that, in this document, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the sentence "comprising a ..." does not exclude the presence of other identical elements in the process, method, article or device comprising the element. In addition, components, features, and elements with the same name in different embodiments of the present application may have the same meaning or different meanings, and their specific meanings need to be determined by their explanation in the specific embodiment or further combined with the context of the specific embodiment.
[0081] It should be understood that although the terms first, second, third, etc. may be used herein to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of this document, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the term "if" as used herein may be interpreted as "when," "when," or "in response to a determination." Furthermore, as used herein, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context indicates otherwise. It should be further understood that the terms "comprising" and "including" indicate the presence of the described features, steps, operations, elements, components, items, types, and / or groups, but do not exclude the presence, occurrence, or addition of one or more other features, steps, operations, elements, components, items, types, and / or groups. The terms "or," "and / or," "including at least one of the following," etc., used herein, may be interpreted as inclusive, or mean any one or any combination. For example, “comprising at least one of the following: A, B, C” means “any of the following: A; B; C; A and B; A and C; B and C; A and B and C”; and for another example, “A, B or C” or “A, B and / or C” means “any of the following: A; B; C; A and B; A and C; B and C; A and B and C”. An exception to this definition will occur only when a combination of elements, functions, steps or operations are inherently mutually exclusive in some manner.
[0082] It should be understood that, although the various steps in the flowchart in the embodiment of the present application are shown in sequence according to the indication of the arrows, these steps are not necessarily performed in sequence in the order indicated by the arrows. Unless clearly stated herein, the execution of these steps is not strictly limited in order, and they can be performed in other orders. Moreover, at least a portion of the steps in the figure may include multiple sub-steps or multiple stages, and these sub-steps or stages are not necessarily performed at the same time, but can be performed at different times, and their execution order is not necessarily performed in sequence, but can be performed in turn or alternately with at least a portion of other steps or sub-steps or stages of other steps.
[0083] As used herein, the words "if" and "if" may be interpreted as "at the time of" or "when" or "in response to determining" or "in response to detecting," depending on the context. Similarly, the phrases "if it is determined" or "if (stated condition or event) is detected" may be interpreted as "when it is determined" or "in response to the determination" or "when detecting (stated condition or event)" or "in response to detecting (stated condition or event)," depending on the context.
[0084] It should be noted that in this article, step codes such as S1 are used for the purpose of expressing the corresponding content more clearly and concisely, and do not constitute a substantial limitation on the sequence.
[0085] It should be understood that the specific embodiments described herein are only used to explain the present application and are not intended to limit the present application.
[0086] In the subsequent description, the use of suffixes such as "module", "component" or "unit" to represent elements is only for the purpose of facilitating the description of the present application and has no specific meaning. Therefore, "module", "component" or "unit" can be used interchangeably.
[0087] The processing device in this application can be a smart terminal or a server, etc., and the smart terminal can be implemented in various forms. For example, the smart terminal can include smart terminals such as mobile phones, tablet computers, laptop computers, PDAs, portable media players (PMPs), navigation devices, wearable devices, smart bracelets, pedometers, etc., as well as fixed terminals such as digital TVs and desktop computers.
[0088] The subsequent description will be made by taking a mobile terminal as an example. It will be understood by those skilled in the art that, in addition to components specifically used for mobile purposes, the configuration according to the embodiments of the present application can also be applied to fixed-type terminals.
[0089] See also Figure 1 , which is a schematic diagram of the hardware structure of a mobile terminal for implementing various embodiments of the present application. The mobile terminal 100 may include: an RF (Radio Frequency) unit 101, a WiFi module 102, an audio output unit 103, an A / V (audio / video) input unit 104, a sensor 105, a display unit 106, a user input unit 107, an interface unit 108, a memory 109, a processor 110, and a power supply 111. Those skilled in the art will understand that Figure 1 The structure of the mobile terminal shown in the figure does not constitute a limitation to the mobile terminal. The mobile terminal may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently.
[0090] The following combination Figure 1 A detailed introduction to the various components of the mobile terminal:
[0091] The RF unit 101 can be used to send and receive information or receive signals during calls. Specifically, it receives downlink information from the base station and transmits it to the processor 110 for processing. It also transmits uplink data to the base station. Typically, the RF unit 101 includes, but is not limited to, an antenna, at least one amplifier, a transceiver, a coupler, a low-noise amplifier, a duplexer, and more. Furthermore, the RF unit 101 can communicate with the network and other devices via wireless communication. The above-mentioned wireless communications may use any communication standard or protocol, including but not limited to GSM (Global System of Mobile communication), GPRS (General Packet Radio Service), CDMA2000 (Code Division Multiple Access 2000), WCDMA (Wideband Code Division Multiple Access), TD-SCDMA (Time Division-Synchronous Code Division Multiple Access), FDD-LTE (Frequency Division Duplexing-Long Term Evolution), TDD-LTE (Time Division Duplexing-Long Term Evolution), 5G and 6G, etc.
[0092] WiFi is a short-range wireless transmission technology. Mobile terminals can help users send and receive emails, browse web pages, and access streaming media through the WiFi module 102. It provides users with wireless broadband Internet access. Figure 1 The WiFi module 102 is shown, but it is understandable that it is not an essential component of the mobile terminal and can be omitted as needed without changing the essence of the invention.
[0093] The audio output unit 103 can convert audio data received by the RF unit 101 or the WiFi module 102 or stored in the memory 109 into an audio signal and output it as sound when the mobile terminal 100 is in a call signal reception mode, a talk mode, a recording mode, a voice recognition mode, a broadcast reception mode, or the like. Furthermore, the audio output unit 103 can also provide audio output related to a specific function performed by the mobile terminal 100 (e.g., a call signal reception sound, a message reception sound, etc.). The audio output unit 103 may include a speaker, a buzzer, or the like.
[0094] The A / V input unit 104 is used to receive audio or video signals. The A / V input unit 104 may include a graphics processing unit (GPU) 1041 and a microphone 1042. The GPU 1041 processes image data of still images or videos captured by an image capture device (e.g., a camera) in video capture mode or image capture mode. The processed image frames may be displayed on the display unit 106. The image frames processed by the GPU 1041 may be stored in the memory 109 (or other storage medium) or transmitted via the RF unit 101 or the WiFi module 102. The microphone 1042 may receive sound (audio data) in operating modes such as a phone call mode, a recording mode, and a voice recognition mode, and may process such sound into audio data. In the phone call mode, the processed audio (voice) data may be converted into a format that can be transmitted to a mobile communication base station via the RF unit 101. The microphone 1042 may implement various types of noise cancellation (or suppression) algorithms to eliminate (or suppress) noise or interference generated during the reception and transmission of audio signals.
[0095] The mobile terminal 100 also includes at least one sensor 105, such as a light sensor, a motion sensor, and other sensors. Optionally, the light sensor includes an ambient light sensor and a proximity sensor. Optionally, the ambient light sensor can adjust the brightness of the display panel 1061 according to the brightness of the ambient light, and the proximity sensor can turn off the display panel 1061 and / or the backlight when the mobile terminal 100 is moved to the ear. As a type of motion sensor, the accelerometer sensor can detect the magnitude of acceleration in all directions (generally three axes), and can detect the magnitude and direction of gravity when stationary. It can be used for applications that recognize the posture of the mobile phone (such as horizontal and vertical screen switching, related games, magnetometer posture calibration), vibration recognition related functions (such as pedometer, tapping), etc.; as for other sensors that can be configured in the mobile phone, such as fingerprint sensors, pressure sensors, iris sensors, molecular sensors, gyroscopes, barometers, hygrometers, thermometers, infrared sensors, etc., they will not be described here.
[0096] The display unit 106 is used to display information input by the user or information provided to the user. The display unit 106 may include a display panel 1061, which may be configured in the form of a liquid crystal display (LCD), an organic light-emitting diode (OLED), or the like.
[0097] The user input unit 107 can be used to receive input digital or character information, and to generate key signal input related to the user settings and function control of the mobile terminal. Optionally, the user input unit 107 may include a touch panel 1071 and other input devices 1072. The touch panel 1071, also known as a touch screen, can collect user touch operations on or near it (such as operations performed by the user using a finger, stylus, or any other suitable object or accessory on or near the touch panel 1071) and drive the corresponding connection device according to a pre-set program. The touch panel 1071 may include two parts: a touch detection device and a touch controller. Optionally, the touch detection device detects the user's touch direction and detects the signal generated by the touch operation, and transmits the signal to the touch controller; the touch controller receives the touch information from the touch detection device and converts it into touch point coordinates, which are then sent to the processor 110. It can also receive commands sent by the processor 110 and execute them. In addition, the touch panel 1071 can be implemented using various types such as resistive, capacitive, infrared, and surface acoustic wave. In addition to the touch panel 1071, the user input unit 107 may further include other input devices 1072. Optionally, the other input devices 1072 may include, but are not limited to, one or more of a physical keyboard, function keys (such as volume control keys, power keys, etc.), a trackball, a mouse, a joystick, etc., and the specifics are not limited here.
[0098] Optionally, the touch panel 1071 may cover the display panel 1061. When the touch panel 1071 detects a touch operation on or near it, it transmits the information to the processor 110 to determine the type of touch event. The processor 110 then provides a corresponding visual output on the display panel 1061 according to the type of touch event. Figure 1 In the embodiment, the touch panel 1071 and the display panel 1061 are two independent components to realize the input and output functions of the mobile terminal. However, in some embodiments, the touch panel 1071 and the display panel 1061 can be integrated to realize the input and output functions of the mobile terminal, which is not limited here.
[0099] The interface unit 108 serves as an interface through which at least one external device can be connected to the mobile terminal 100. For example, the external device may include a wired or wireless headset port, an external power supply (or battery charger) port, a wired or wireless data port, a memory card port, a port for connecting a device with an identification module, an audio input / output (I / O) port, a video I / O port, a headphone port, etc. The interface unit 108 may be used to receive input (e.g., data information, power, etc.) from an external device and transmit the received input to one or more elements within the mobile terminal 100 or may be used to transmit data between the mobile terminal 100 and an external device.
[0100] Memory 109 can be used to store software programs and various data. Memory 109 may primarily include a program storage area and a data storage area. Optionally, the program storage area may store an operating system and at least one application required for a function (such as a sound playback function or an image playback function); the data storage area may store data generated based on the use of the mobile phone (such as audio data, a phone book, etc.). Furthermore, memory 109 may include high-speed random access memory and non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device.
[0101] Processor 110 is the control center of the mobile terminal, connecting all components of the mobile terminal using various interfaces and circuits. By running or executing software programs and / or modules stored in memory 109 and accessing data stored in memory 109, it executes various functions of the mobile terminal and processes data, thereby providing overall monitoring of the mobile terminal. Processor 110 may include one or more processing units; preferably, processor 110 may integrate an application processor and a modem processor. Optionally, the application processor primarily handles the operating system, user interface, and application programs, while the modem processor primarily handles wireless communications. It is understood that the modem processor may not be integrated into processor 110.
[0102] The mobile terminal 100 may also include a power supply 111 (such as a battery) for supplying power to various components. Preferably, the power supply 111 may be logically connected to the processor 110 through a power management system, thereby managing functions such as charging, discharging, and power consumption through the power management system.
[0103] although Figure 1 Not shown, the mobile terminal 100 may further include a Bluetooth module, etc., which will not be described in detail here.
[0104] To facilitate understanding of the embodiments of the present application, the communication network system on which the mobile terminal of the present application is based is described below.
[0105] See also Figure 2 , Figure 2 A communication network system architecture diagram is provided for an embodiment of the present application. The communication network system is an LTE system of universal mobile communication technology. The LTE system includes a UE (User Equipment) 201, an E-UTRAN (Evolved UMTS Terrestrial Radio Access Network) 202, an EPC (Evolved Packet Core) 203 and an operator's IP service 204, which are connected in sequence.
[0106] Optionally, UE201 may be the above-mentioned terminal 100, which will not be described in detail here.
[0107] E-UTRAN 202 includes eNodeB 2021 and other eNodeBs 2022 . Optionally, eNodeB 2021 may be connected to other eNodeBs 2022 via a backhaul (eg, an X2 interface). eNodeB 2021 is connected to EPC 203 , and eNodeB 2021 may provide access from UE 201 to EPC 203 .
[0108] EPC 203 may include an MME (Mobility Management Entity) 2031, an HSS (Home Subscriber Server) 2032, other MMEs 2033, an SGW (Serving Gate Way) 2034, a PGW (PDN Gate Way) 2035, and a PCRF (Policy and Charging Rules Function) 2036. Optionally, MME 2031 is a control node that processes signaling between UE 201 and EPC 203, providing bearer and connection management. HSS 2032 provides registers for managing functions such as the Home Location Register (not shown) and stores user-specific information such as service features and data rates. All user data can be sent through SGW2034, PGW2035 can provide IP address allocation and other functions for UE 201, PCRF2036 is the policy and charging control policy decision point for service data flow and IP bearer resources, and it selects and provides available policy and charging control decisions for the policy and charging execution function unit (not shown in the figure).
[0109] The IP service 204 may include the Internet, an intranet, an IMS (IP Multimedia Subsystem), or other IP services.
[0110] Although the above introduction takes the LTE system as an example, those skilled in the art should know that this application is not only applicable to the LTE system, but can also be applied to other wireless communication systems, such as GSM, CDMA2000, WCDMA, TD-SCDMA, 5G and future new network systems (such as 6G), etc., which are not limited here.
[0111] Figure 3 This is a schematic diagram of the hardware structure of a controller 140 provided in this application. The controller 140 includes a memory 1401 and a processor 1402. The memory 1401 is used to store program instructions, and the processor 1402 is used to call the program instructions in the memory 1401 to execute the steps performed by the controller in the first embodiment of the above method. The implementation principles and beneficial effects are similar and will not be repeated here.
[0112] Optionally, the controller further includes a communication interface 1403, which can be connected to the processor 1402 via a bus 1404. The processor 1402 can control the communication interface 1403 to implement the receiving and sending functions of the controller 140.
[0113] Figure 4 This is a schematic diagram of the hardware structure of a network node 150 provided in this application. The network node 150 includes a memory 1501 and a processor 1502. The memory 1501 is used to store program instructions, and the processor 1502 is used to call the program instructions in the memory 1501 to execute the steps performed by the first node in the first embodiment of the above method. The implementation principles and beneficial effects are similar and will not be repeated here.
[0114] Optionally, the controller further includes a communication interface 1503, which can be connected to the processor 1502 via a bus 1504. The processor 1502 can control the communication interface 1503 to implement the receiving and sending functions of the network node 150.
[0115] The integrated modules implemented in the form of software function modules can be stored in a computer-readable storage medium. The software function modules stored in a storage medium include a number of instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) or a processor to execute some of the steps of the methods of various embodiments of the present application.
[0116] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When software is used for implementation, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function according to the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a storage medium or transmitted from one storage medium to another storage medium. For example, the computer instructions can be transmitted from a website, computer, server or data center to another website, computer, server or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) method. The storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more available media integrations. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid state drive solid state disk, SSD), etc.
[0117] Based on the above-mentioned mobile terminal hardware structure and communication network system, various embodiments of the present application are proposed.
[0118] First embodiment
[0119] Reference Figure 5 , Figure 5 FIG. 1 is a flow chart of an image processing method according to a first embodiment. The image processing method according to the embodiment of the present application can be applied to a processing device, comprising step S1:
[0120] Step S1 : performing prediction processing on at least one image block according to the spatial domain features of at least one reference block.
[0121] In this embodiment, the processing device can be a smart terminal, such as a mobile phone, a computer, etc., or a server, such as a local server or a cloud server. In this embodiment and this application, the processing device is mainly described as a smart terminal.
[0122] Optionally, the technical solution of this embodiment can be applied to the fields of image coding and decoding, video coding and decoding, hardware video coding and decoding, dedicated circuit video coding and decoding, real-time video coding and decoding, and so on.
[0123] Optionally, the processing device may store various images and videos in advance and select an image to be predicted from each image as an image block, or segment the selected image and use the segmented image block as the image block to be predicted. Alternatively, a frame of image may be extracted from a video sequence as an image block, or the extracted frame of image may be segmented to obtain an image block. Alternatively, the processing device may receive an input image or video and extract a frame of image from the image or video as an image block, or segment the extracted frame of image to obtain an image block. Alternatively, the processing device may receive an image or video sent from another network device and extract a frame of image from the image or video as an image block, or segment the extracted frame of image to obtain an image block. In this case, the processing device may pre-establish a communication connection with a network device on the network side of the mobile communication system in which it is located. Thus, the network device can transmit the image or video to the terminal device via the communication connection, and the terminal device will then receive the image or video.
[0124] Optionally, the technical solution of this embodiment can be performed for intra-frame prediction or inter-frame prediction, without limitation herein.
[0125] Optionally, the spatial domain feature of the reference block may include at least one of a style feature, a gradient distribution, an edge direction histogram, a color component correlation, and a motion activity of the reference block.
[0126] Optionally, the gradient distribution is such as the gradient calculation of a VVC (Versatile Video Coding, a new generation video coding standard jointly developed by the International Telecommunication Union Telecommunication Standardization Sector and the International Organization for Standardization / International Electrotechnical Commission Moving Picture Experts Group) ALF (Adaptive Loop Filter).
[0127] Optionally, the edge direction histogram corresponds to AV1 (AOMedia Video 1, an open source, royalty-free, next-generation video coding standard developed by the Alliance for Open Media) directional filtering.
[0128] Optionally, color component correlation is overlaid on VVC CCLM (Cross-Component Linear Model) chrominance prediction. Optionally, motion activity is used for dynamic QP (quantization parameter) adjustment.
[0129] Alternatively, one reference block may correspond to one spatial feature, or two or more reference blocks may correspond to one spatial feature, or one reference block may correspond to two or more spatial features, or two or more reference blocks may correspond to two or more spatial features. Alternatively, each reference pixel in a reference block may correspond to at least one spatial feature.
[0130] Optionally, a spatial feature may be determined or obtained based on at least one reference pixel in at least one reference block.
[0131] Optionally, the spatial feature of the reference block may include at least one spatial feature determined or obtained based on at least one reference pixel or a first reference intermediate element in the reference block.
[0132] Optionally, a spatial feature can be determined or obtained based on a reference pixel or a first reference intermediate element, or two or more spatial features can be determined or obtained based on a reference pixel or a first reference intermediate element, or a spatial feature can be determined or obtained based on two or more reference pixels or first reference intermediate elements, or two or more spatial features can be determined or obtained based on two or more reference pixels or first reference intermediate elements.
[0133] Optionally, the spatial features of the reference block may include the spatial features corresponding to all or part of the reference pixels or first reference intermediate elements in the reference block. The spatial features corresponding to a reference pixel or a first reference intermediate element can be determined or obtained based on at least one reference pixel or first reference intermediate element in the reference block, such as determined or obtained based on the reference pixel or first reference intermediate element and other reference pixels or first reference intermediate elements in the reference block.
[0134] Optionally, the spatial domain feature of the at least one reference block may be determined or obtained based on a statistical feature value of at least one reference pixel or a first reference intermediate element in the at least one reference block.
[0135] Optionally, the spatial domain features of at least one reference block may be determined or obtained by preprocessing the reference block and then determining or obtaining the statistical feature value of at least one reference pixel or first reference intermediate element / element in the preprocessed reference block.
[0136] Optionally, the statistical characteristic value may include at least one of the mean, maximum, minimum, median, mode (the pixel value with the highest frequency), range (the difference between the maximum and minimum values), variance, standard deviation, etc., or may include values determined or obtained according to other rules or calculation methods, which are not limited here.
[0137] Optionally, the spatial domain features of the reference block may be determined or obtained by sequentially calculating statistical feature values of all reference pixels or first reference intermediate elements in the reference block, and based on these statistical feature values.
[0138] Optionally, the reference block is an image block that has been predicted and / or reconstructed.
[0139] Optionally, the reference block of the at least one image block may be another image block used to assist in performing prediction processing on the at least one image block.
[0140] Optionally, the reference block may be a rectangular block, a non-rectangular block, an L-shaped block, or a T-shaped block.
[0141] Optionally, the reference block may include multiple rows and / or multiple columns of pixels, or may include one row and / or one column of pixels.
[0142] Optionally, the reference block of at least one image block includes a reference prediction block and / or a reference reconstruction block of the reference block.
[0143] Optionally, the reference block is determined or obtained according to at least one of the following methods 1 to 6:
[0144] Mode 1: at least one of the upper adjacent pixel, upper non-adjacent pixel, left adjacent pixel, left non-adjacent pixel, upper left adjacent pixel, and upper left non-adjacent pixel of the image block;
[0145] Optionally, the pixel to be predicted in the image block may be used as the first pixel.
[0146] Optionally, the upper adjacent pixel may be a pixel in the same image that is located in the same image block obtained by final division together with the first pixel, and / or the pixel is located above the first pixel and adjacent to the first pixel.
[0147] Optionally, the upper non-adjacent pixel may be a pixel in the same image that is in the same image block obtained by finally dividing the same image as the first pixel. Although the pixel is located above the first pixel, the pixel is not adjacent to the first pixel.
[0148] Optionally, the left adjacent pixel may be a pixel in the same image that is located in the same image block obtained by final division together with the first pixel, and / or the pixel is located on the left side of the first pixel and adjacent to the first pixel.
[0149] Optionally, the left non-adjacent pixel may be a pixel in the same image that is in the same image block obtained by the final division as the first pixel. Although the pixel is located to the left of the first pixel, the pixel is not adjacent to the first pixel.
[0150] Optionally, the upper left adjacent pixel may be a pixel in the same image that is in the same image block obtained by final division together with the first pixel, and / or the pixel is located above and to the left of the first pixel and adjacent to the first pixel.
[0151] Optionally, the upper left non-adjacent pixel may be a pixel in the same image that is in the same image block obtained by the final division together with the first pixel. Although the pixel is located to the upper left of the first pixel, the pixel is not adjacent to the first pixel.
[0152] Optionally, the lower left adjacent pixel may be a pixel in the same image that is in the same image block obtained by final division as the first pixel, and / or the pixel is located below the left of the first pixel and adjacent to the first pixel.
[0153] Optionally, the lower left non-adjacent pixel may be a pixel in the same image that is in the same image block obtained by final division as the first pixel, and / or the pixel is located below the left of the first pixel and is not adjacent to the first pixel.
[0154] Optionally, the upper right adjacent pixel may be a pixel in the same image that is in the same image block obtained by final division as the first pixel, and / or the pixel is located above and to the right of the first pixel and adjacent to the first pixel.
[0155] Optionally, the upper right non-adjacent pixel may be a pixel in the same image that is in the same image block obtained by final division as the first pixel, and / or the pixel is located above and to the right of the first pixel and is not adjacent to the first pixel.
[0156] Optionally, at least one of the upper adjacent pixel, upper non-adjacent pixel, left adjacent pixel, left non-adjacent pixel, upper left adjacent pixel, upper left non-adjacent pixel, lower left adjacent pixel, lower left non-adjacent pixel, upper right adjacent pixel and upper right non-adjacent pixel can be a reconstructed pixel or a predicted pixel.
[0157] Optionally, at least one of the upper adjacent pixel, the upper non-adjacent pixel, the left adjacent pixel, the left non-adjacent pixel, the upper left adjacent pixel, the upper left non-adjacent pixel, the lower left adjacent pixel, the lower left non-adjacent pixel, the upper right adjacent pixel and the upper right non-adjacent pixel can be directly used as a pixel in the reference block, or the at least one acquired pixel can be deduced or calculated to obtain the pixel in the reference block, and / or the reference block can be determined or obtained based on the pixels in the reference block.
[0158] In this embodiment, a reference block is determined or generated based on at least one of the upper adjacent pixels, upper non-adjacent pixels, left adjacent pixels, left non-adjacent pixels, upper left adjacent pixels, upper left non-adjacent pixels, lower left adjacent pixels, lower left non-adjacent pixels, upper right adjacent pixels, and upper right non-adjacent pixels of the image block, and a prediction block is determined or generated based on a derivative block corresponding to at least one reference block. Specifically, when performing prediction processing, the corresponding derivative block can be determined based on an accurate and effective reference block, and prediction processing can be performed using the derivative block, thereby improving the prediction effect of the prediction processing.
[0159] Mode 2: at least one of a neighbor block, a non-neighbor block, a cross-component block, a co-located block, a time domain block, and a default block corresponding to the image block;
[0160] Optionally, the default block may be a block set in advance, for example, a block with typical pixel features pre-set by an encoder and / or a decoder.
[0161] Optionally, the neighbor block may be a block adjacent to the image block, and / or may be a block that has been predicted or reconstructed.
[0162] Optionally, the non-neighbor block may be a block that is not adjacent to the image block, and / or may be a block that has been predicted or reconstructed.
[0163] Optionally, the co-located block may be an image block in the co-located image with the same position and size as the image block. Optionally, the co-located image may be an image in the reference image that is closest to the current image in terms of time.
[0164] Optionally, the time domain block can be a block distinguished from the time domain, such as an image block in the previous frame. For example, if there is video data containing three frames of images, the first frame of image is played in the first second, the second frame of image is played in the second second, and the third frame of image is played in the third second. If the image block predicted at the current moment (such as the block to be predicted) is the image block divided from the second frame of image, then the time domain block can be determined to be the image block corresponding to it in the first frame of image.
[0165] Optionally, the cross-component block may be an image block of a different component than the current at least one image block. For example, if the current at least one image block is an image block of the Y component, the cross-component block may be an image block of the U component and / or the V component. Optionally, if the current at least one image block is an image block of the U component, the cross-component block may be an image block of the Y component and / or the V component. Optionally, if the current at least one image block is an image block of the V component, the cross-component block may be an image block of the U component and / or the Y component.
[0166] Optionally, at least one pixel may be obtained from at least one of a neighboring block, a non-neighboring block, a cross-component block, a co-located block, a time-domain block, and a default block corresponding to the image block. The obtained at least one pixel may be used as a pixel in the reference block, or the obtained at least one pixel may be derived or calculated to obtain the pixel in the reference block. The reference block is determined or obtained based on the pixels in the reference block.
[0167] Optionally, at least one of an upper adjacent pixel, an upper non-adjacent pixel, a left adjacent pixel, a left non-adjacent pixel, an upper left adjacent pixel, and an upper left non-adjacent pixel can be determined from at least one of a neighboring block, a non-neighboring block, a cross-component block, a co-located block, a time-domain block, and a default block corresponding to the image block, and used as a pixel of the reference block. Furthermore, at least one of an adjacent area and / or a non-adjacent area corresponding to the image block can be determined, and at least one pixel therefrom can be selected as a pixel of the reference block. The reference block is determined or obtained based on the pixels in the reference block.
[0168] In this embodiment, a reference block is determined or generated based on at least one of a neighboring block, a non-neighboring block, a cross-component block, a co-located block, a time domain block, and a default block corresponding to the image block, and a prediction block is determined or generated based on a derivative block corresponding to at least one reference block. Specifically, when performing prediction processing, the corresponding derivative block can be determined based on an accurate and effective reference block, and prediction processing can be performed using the derivative block, thereby improving the prediction effect of the prediction processing.
[0169] Mode 3: at least one of the width, height, block size, and block area of the image block;
[0170] Optionally, at least one reference block is determined or obtained according to at least one of width, height, block size and block area of at least one image block.
[0171] Optionally, for example, if at least one of the width, height, block size and block area of at least one image block is greater than a preset threshold, at least one image block is selected as a reference block from the neighbor blocks, non-neighbor blocks, cross-component blocks, co-located blocks, time domain blocks and default blocks corresponding to the at least one image block.
[0172] Optionally, at least one of the width, height, block size and block area of at least one image block may be input into a neural network for determining a reference block, and the reference block may be output.
[0173] In this embodiment, a reference block is determined or generated based on at least one of the width, height, block size, and block area of an image block, and a prediction block is determined or generated based on a derivative block corresponding to at least one reference block. Specifically, when performing prediction processing, the corresponding derivative block can be determined based on an accurate and effective reference block, and the prediction processing is performed using the derivative block, thereby improving the prediction effect of the prediction processing.
[0174] Mode 4, a candidate block determined or generated by a candidate motion vector or a candidate block vector of an image block;
[0175] Optionally, the candidate motion vector or candidate block vector of the image block may include the motion vector or block vector corresponding to at least one of the neighboring blocks, non-neighboring blocks, cross-component blocks, co-located blocks, time domain blocks and default blocks corresponding to the image block, and may also include the motion vector or block vector corresponding to at least one of the upper adjacent pixels, upper non-adjacent pixels, left adjacent pixels, left non-adjacent pixels, upper left adjacent pixels and upper left non-adjacent pixels of the image block, and may also include the motion vector or block vector of the image block, etc. The following only takes the motion vector or block vector of the image block as an example.
[0176] Optionally, a block vector calculation is performed on the image block, and a candidate block corresponding to the image block is determined according to the block vector calculation result, for example, pixels corresponding to the block vector calculation result are used as pixels in the candidate block.
[0177] Optionally, a motion vector is calculated for the image block, and a candidate block corresponding to the image block is determined according to the motion vector calculation result, for example, pixels corresponding to the motion vector calculation result are used as pixels in the candidate block.
[0178] In this embodiment, a reference block is determined or generated by a candidate block determined or generated by a candidate motion vector or a candidate block vector of an image block, and a prediction block is determined or generated based on a derivative block corresponding to at least one reference block. Specifically, when performing prediction processing, the corresponding derivative block can be determined based on an accurate and effective reference block, and the prediction processing is performed using the derivative block, which can improve the prediction effect of the prediction processing.
[0179] Mode 5: If the first information of the image block satisfies the first condition, the reference block is the first reference block;
[0180] Optionally, the first information of the image block may be at least one of the above-mentioned methods 1 to 7.
[0181] Optionally, satisfying the first condition may be any condition set in advance by the user, such as the width of the image block being greater than a preset width threshold, the height of the image block being greater than a preset height threshold, etc.
[0182] Optionally, the first reference block may be a reference block set and determined in advance, such as at least one of a neighbor block, a non-neighbor block, a cross-component block, a co-located block, a time domain block, a candidate block and a default block corresponding to the image block.
[0183] Optionally, when the first information of the image block satisfies the first condition, a first reference block of the reference block may be determined, and a prediction block may be determined or generated based on a derivative block corresponding to the first reference block of at least one image block.
[0184] In this embodiment, when the first information of the image block satisfies the first condition, the reference block is the first reference block, and the prediction block is determined or generated based on the derivative block corresponding to the first reference block. Specifically, when performing the prediction processing, the corresponding derivative block can be determined based on the accurate and effective reference block, and the prediction processing is performed using the derivative block, which can improve the prediction effect of the prediction processing.
[0185] Mode 6: If the first information of the image block satisfies the second condition, the reference block is the second reference block.
[0186] Optionally, satisfying the second condition may be any condition set in advance by the user, and / or may be different from the second condition. For example, if satisfying the first condition is that the width of the image block is greater than a preset width threshold, satisfying the second condition may be that the width of the image block is less than the preset width threshold.
[0187] Optionally, the second reference block may be a reference block set and determined in advance, such as a block other than the first reference block, which may be a neighbor block, a non-neighbor block, a cross-component block, a co-located block, a time domain block, a candidate block, and a default block corresponding to the image block.
[0188] In this embodiment, when the first information of the image block satisfies the second condition, the reference block is the second reference block, and the prediction block is determined or generated based on the derivative block corresponding to the second reference block. Then, when performing the prediction processing, the corresponding derivative block can be determined based on the accurate and effective reference block, and the derivative block can be used to perform the prediction processing, which can improve the prediction effect of the prediction processing.
[0189] Optionally, in step S1 , prediction processing may be performed on at least one image block according to the spatial domain features of at least one reference block to determine or obtain a prediction block.
[0190] Optionally, the prediction block may be an image block that has undergone prediction processing.
[0191] Optionally, the prediction block may include at least one predicted pixel, and may also include pixels associated with the at least one predicted pixel.
[0192] Optionally, the predicted pixel may be a pixel predicted at the encoding side, which may be referred to as a predicted pixel, and the reconstructed pixel may be a pixel predicted at the decoding side, which may be referred to as a predicted pixel.
[0193] Optionally, the processing device may be a decoding end. If at the decoding end, the prediction block may be a decoded image block, and / or the prediction block may include reconstruction of pixels.
[0194] Optionally, the processing device may be an encoding end. If at the encoding end, the prediction block may be an image block that has undergone prediction processing, and / or the prediction block may include pixel predictions.
[0195] In this embodiment, when performing prediction processing on at least one image block, spatial characteristics of a reference block of at least one image block are comprehensively considered, thereby improving the prediction effect of the prediction processing.
[0196] Second embodiment
[0197] Based on the first embodiment, a second embodiment is proposed.
[0198] In this embodiment, the spatial feature is determined or obtained according to at least one of the following methods 1 to 8:
[0199] Method 1: statistical feature values of a reference pixel and at least one neighboring pixel in at least one reference block;
[0200] Optionally, the statistical characteristic value can be an indicator that quantitatively describes the statistical properties of pixels, which can be the characteristic mean (such as the mean of pixel values), maximum value, minimum value, or variance, mean square error, mode, median, and is not limited here.
[0201] Optionally, the neighbor pixel may be a reference pixel that is adjacent to or in contact with the reference pixel within a certain range.
[0202] Optionally, the reference pixel may be a selected pixel in a reference block.
[0203] Optionally, the statistical characteristic values of the reference pixel and at least one neighboring pixel in at least one reference block can be statistically determined, and the spatial domain feature can be determined or obtained based on the at least one statistical characteristic value. For example, the correspondence between each characteristic value and the spatial domain feature can be set in advance, and the spatial domain feature can be determined or obtained based on the correspondence.
[0204] Optionally, at least one spatial feature may be determined or obtained based on statistical feature values of a reference pixel and at least one neighboring pixel in at least one reference block, and prediction processing may be performed on at least one image block based on the at least one spatial feature.
[0205] In this embodiment, spatial features are determined or obtained based on statistical characteristic values of reference pixels and at least one neighboring pixel in at least one reference block, thereby ensuring the effectiveness of the spatial features. When using the spatial features to perform prediction processing on at least one image block, key image block features can be focused, thereby improving the prediction effect of at least one image block.
[0206] Method 2: statistical characteristic values of at least two non-neighboring pixels in at least one reference block;
[0207] Optionally, the non-neighbor pixel may be a pixel in the reference block that has a relatively long pixel distance from the reference pixel, and the pixel distance between the two is greater than a pixel distance threshold, and the pixel distance threshold may be set in advance.
[0208] Optionally, the statistical characteristic values of at least two non-neighbor pixels in at least one reference block can be calculated and counted to determine or obtain the corresponding statistical characteristic values, and then the spatial domain features can be determined or obtained based on the statistical characteristic values of at least two non-neighbor pixels in at least one reference block. For example, the correspondence between each characteristic value and the spatial domain feature can be set in advance, and the spatial domain feature can be determined or obtained based on the correspondence.
[0209] Optionally, at least one spatial feature may be determined or obtained based on statistical feature values of a reference pixel and at least one non-neighboring pixel in at least one reference block, and prediction processing may be performed on at least one image block based on the at least one spatial feature.
[0210] In this embodiment, spatial features are determined or obtained based on statistical characteristic values of reference pixels and at least one non-neighboring pixel in at least one reference block, thereby ensuring the effectiveness of the spatial features. When using the spatial features to perform prediction processing on at least one image block, key image block features can be focused, thereby improving the prediction effect on at least one image block.
[0211] Method three: a statistical feature value of a reference pixel in at least one reference block and at least one of the following: at least one upper neighboring pixel of the reference pixel, at least one left neighboring pixel of the reference pixel, at least one right neighboring pixel of the reference pixel, and at least one lower neighboring pixel of the reference pixel;
[0212] Optionally, the adjacent pixels may be pixels adjacent to the reference pixel or the first reference intermediate element in the reference block, and the adjacent pixels of the reference pixel or the first reference intermediate element may include at least one upper adjacent pixel, at least one lower adjacent pixel, at least one right adjacent pixel, and at least one lower adjacent pixel.
[0213] Optionally, the spatial characteristics of at least one reference block can be determined or obtained based on the statistical characteristic values of the reference pixel or the first reference intermediate element in at least one reference block and at least one upper adjacent pixel of the reference pixel or the first reference intermediate element, at least one left adjacent pixel of the reference pixel or the first reference intermediate element, at least one right adjacent pixel of the reference pixel or the first reference intermediate element, and at least one lower adjacent pixel of the reference pixel or the first reference intermediate element, and prediction processing can be performed on at least one image block based on the spatial characteristics of the at least one reference block.
[0214] Optionally, for all or part of the reference pixels or first reference intermediate elements in at least one reference block, statistical characteristic values of the reference pixel or first reference intermediate element in at least one reference block and at least one of the above adjacent pixels of the reference pixel or first reference intermediate element, at least one left adjacent pixel of the reference pixel or first reference intermediate element, at least one right adjacent pixel of the reference pixel or first reference intermediate element, and at least one below adjacent pixel of the reference pixel or first reference intermediate element can be determined, and at least one spatial feature of the at least one reference block can be determined or obtained based on the statistical characteristic values.
[0215] Optionally, the spatial characteristics of at least one reference block can be determined or obtained based on the statistical characteristic values of the reference pixel or the first reference intermediate element in at least one reference block and an upper adjacent pixel, a left adjacent pixel, a right adjacent pixel and a lower adjacent pixel of the reference pixel or the first reference intermediate element, and prediction processing can be performed on at least one image block based on the spatial characteristics of the at least one reference block.
[0216] In this embodiment, spatial features are determined or obtained based on the statistical characteristic values of a reference pixel in at least one reference block and at least one upper adjacent pixel of the reference pixel, at least one left adjacent pixel of the reference pixel, at least one right adjacent pixel of the reference pixel, and at least one lower adjacent pixel of the reference pixel, thereby ensuring the effectiveness of the spatial features. When the spatial features are used to perform prediction processing on at least one image block, key image block features can be focused, thereby improving the prediction effect on at least one image block.
[0217] Method 4: weighted statistical feature values of a reference pixel and pixels adjacent to the reference pixel in at least one reference block;
[0218] Optionally, the weighted statistical feature value may include at least one of the weighted mean, weighted maximum, weighted minimum, weighted median, weighted mode (the pixel value with the highest frequency), weighted range (the difference between the maximum and minimum values), weighted variance, weighted standard deviation, etc. of the reference pixel or the first reference intermediate element in at least one reference block, and the pixels adjacent to the reference pixel or the first reference intermediate element, or may include values determined or obtained according to other rules or calculation methods, which are not limited here.
[0219] Optionally, the spatial characteristics of at least one reference block can be determined or obtained based on the weighted statistical characteristic values of the reference pixel or the first reference intermediate element in at least one reference block and at least one upper adjacent pixel of the reference pixel or the first reference intermediate element, at least one left adjacent pixel of the reference pixel or the first reference intermediate element, at least one right adjacent pixel of the reference pixel or the first reference intermediate element, and at least one lower adjacent pixel of the reference pixel or the first reference intermediate element, and prediction processing can be performed on at least one image block based on the spatial characteristics of the at least one reference block.
[0220] Optionally, the weights of the weighting can be preset.
[0221] Optionally, the weighted weights can be obtained through pre- / offline neural network training.
[0222] Optionally, the weights of different pixels may be the same or different.
[0223] Optionally, for all or part of the reference pixels or first reference intermediate elements in at least one reference block, weighted statistical characteristic values of the reference pixel or first reference intermediate element in at least one reference block and at least one of the above adjacent pixels of the reference pixel or first reference intermediate element, at least one left adjacent pixel of the reference pixel or first reference intermediate element, at least one right adjacent pixel of the reference pixel or first reference intermediate element, and at least one below adjacent pixel of the reference pixel or first reference intermediate element can be determined, and at least one spatial feature of the at least one reference block can be determined or obtained based on the weighted statistical characteristic values.
[0224] Optionally, the spatial characteristics of at least one reference block can be determined or obtained based on the weighted statistical characteristic values of the reference pixel or the first reference intermediate element in at least one reference block and an upper adjacent pixel, a left adjacent pixel, a right adjacent pixel and a lower adjacent pixel of the reference pixel or the first reference intermediate element, and prediction processing can be performed on at least one image block based on the spatial characteristics of the at least one reference block.
[0225] In this embodiment, spatial features are determined or obtained based on weighted statistical feature values of reference pixels and pixels adjacent to the reference pixels in at least one reference block, thereby ensuring the effectiveness of the spatial features. When using the spatial features to perform prediction processing on at least one image block, key image block features can be focused, thereby improving the prediction effect on at least one image block.
[0226] Method 5: Statistical characteristic value of at least one pixel in the outermost layer of a window centered on the reference pixel, where the size parameter of the window is less than or equal to the size parameter of the reference block;
[0227] Optionally, the window may be a window centered on the reference pixel or the first reference intermediate element, and a size parameter of the window may be set to be smaller than or equal to a size parameter of the reference block.
[0228] Optionally, the spatial characteristics of at least one reference block can be determined or obtained based on the statistical characteristic value of at least one pixel in the outermost layer of a window centered on a reference pixel or a first reference intermediate element in at least one reference block, and prediction processing can be performed on at least one image block based on the spatial characteristics of at least one reference block.
[0229] Optionally, for all or part of the reference pixels or the first reference intermediate element in at least one reference block, the statistical characteristic value of at least one pixel in the outermost layer of the window centered on the reference pixel or the first reference intermediate element in at least one reference block can be determined, and at least one spatial domain feature of the at least one reference block can be determined or obtained based on the statistical characteristic value.
[0230] For example, the size parameter of the reference block is 7×7, and the size parameter of the window can be set to 5×5. For the middlemost reference pixel or the first reference intermediate element of the reference block, the spatial domain characteristics of at least one reference block can be determined or obtained based on the statistical characteristic value of at least one pixel in the outermost 16 pixels of the 5×5 window centered on the reference pixel or the first reference intermediate element.
[0231] Optionally, the spatial characteristics of at least one reference block can be determined or obtained based on the weighted statistical characteristic value of at least one pixel in the outermost layer of a window centered on a reference pixel or a first reference intermediate element in at least one reference block, and prediction processing can be performed on at least one image block based on the spatial characteristics of the at least one reference block.
[0232] In this embodiment, the spatial domain feature is determined or obtained based on the statistical feature value of at least one pixel in the outermost layer of a window centered on a reference pixel, and the size parameter of the window is less than or equal to the size parameter of the reference block, thereby ensuring the validity of the spatial domain feature. When the spatial domain feature is used to perform prediction processing on at least one image block, the key image block features can be focused, and the prediction effect on the at least one image block can be improved.
[0233] Method 6: Statistical eigenvalues of eight pixels that are not adjacent to the reference pixel, and the eight pixels are not adjacent to each other;
[0234] Optionally, eight pixels can be determined from pixels in at least one reference block that are not adjacent to the reference pixel or the first reference intermediate element, and the spatial characteristics of the at least one reference block can be determined or obtained based on the statistical characteristic values of the determined eight pixels, and prediction processing can be performed on at least one image block based on the spatial characteristics of the at least one reference block.
[0235] Optionally, for all or part of the reference pixels or the first reference intermediate element in at least one reference block, a statistical characteristic value that is not adjacent to the reference pixel or the first reference intermediate element in at least one reference block can be determined, and at least one spatial domain feature of the at least one reference block can be determined or obtained based on the statistical characteristic value.
[0236] Optionally, for example, the size of the reference block is 5×5. For the centermost reference pixel or the first reference intermediate element in the reference block, 8 pixels can be determined from 16 pixels that are not adjacent to the reference pixel or the first reference intermediate element. Based on the statistical characteristic values of the determined 8 pixels, at least one spatial domain feature of at least one reference block can be determined or obtained.
[0237] Optionally, eight pixels can be determined from pixels in at least one reference block that are not adjacent to the reference pixel or the first reference intermediate element, and the spatial characteristics of at least one reference block can be determined or obtained based on the weighted statistical characteristic values of the determined eight pixels. At least one image block is predicted based on the spatial characteristics of the at least one reference block, and the weighted weights of each pixel can be the same or different.
[0238] Optionally, the eight pixels that are not adjacent to the reference pixel or the first reference intermediate element may not be adjacent to each other. For example, if the size of the reference block is 5×5, for the centermost reference pixel or the first reference intermediate element in the reference block, eight non-adjacent pixels may be determined from the 16 pixels that are not adjacent to the reference pixel or the first reference intermediate element. Based on the statistical feature values of the determined eight non-adjacent pixels, at least one spatial feature of the at least one reference block may be determined or obtained, and prediction processing may be performed on the at least one image block based on the spatial feature of the at least one reference block.
[0239] In this embodiment, the spatial domain features are determined or obtained based on the statistical feature values of eight pixels that are not adjacent to the reference pixel, and the features that the eight pixels are not adjacent to each other, thereby ensuring the effectiveness of the spatial domain features. When the spatial domain features are used to predict at least one image block, the key image block features can be focused, and the prediction effect of at least one image block can be improved.
[0240] Method seven, statistical feature values of a reference sub-block including at least one of at least one reference pixel, at least one neighboring pixel, and at least one non-neighboring pixel in at least one reference block;
[0241] Optionally, the reference block may include at least one reference sub-block, and the reference sub-block may include at least one of at least one reference pixel, at least one neighboring pixel, and at least one non-neighboring pixel.
[0242] Optionally, each reference sub-block in a reference block may correspond to at least one spatial domain feature.
[0243] Optionally, a spatial feature of the reference block may be determined or obtained based on at least one reference sub-block in at least one reference block. Optionally, the spatial feature of the reference block may include at least one spatial feature determined or obtained based on at least one reference sub-block in the reference block.
[0244] Optionally, the reference sub-block may include at least one reference pixel or a first reference intermediate element and at least one neighboring pixel of the reference pixel or the first reference intermediate element, or the reference sub-block may include at least one reference pixel or a first reference intermediate element and at least one non-neighboring pixel of the reference pixel or the first reference intermediate element, or the reference sub-block may include at least one reference pixel or a first reference intermediate element, at least one neighboring pixel of the reference pixel or the first reference intermediate element, and at least one non-neighboring pixel of the reference pixel or the first reference intermediate element.
[0245] Optionally, a reference sub-block including at least one reference pixel or a first reference intermediate element, at least one neighboring pixel and at least one non-neighboring pixel can be determined from the reference block, and the spatial domain features of at least one reference block can be determined or obtained based on the statistical feature values of the reference sub-block, and prediction processing can be performed on at least one image block based on the spatial domain features of the at least one reference block.
[0246] Optionally, for all or part of the reference pixels or first reference intermediate elements in at least one reference block, a reference sub-block in at least one reference block containing at least one of the reference pixel or the first reference intermediate element, at least one neighboring pixel of the reference pixel or the first reference intermediate element, and at least one non-neighboring pixel of the reference pixel or the first reference intermediate element can be determined, and the spatial domain characteristics of the at least one reference block can be determined or obtained based on the statistical characteristic values of the reference sub-block, and prediction processing can be performed on at least one image based on the spatial domain characteristics of the at least one reference block.
[0247] Optionally, a reference sub-block including at least one reference pixel or a first reference intermediate element, at least one neighboring pixel and at least one non-neighboring pixel can be determined from the reference block, and the spatial domain features of at least one reference block can be determined or obtained based on the weighted statistical feature values of the reference sub-block, and prediction processing can be performed on at least one image block based on the spatial domain features of the at least one reference block.
[0248] In this embodiment, spatial features are determined or obtained based on statistical feature values of a reference sub-block containing at least one reference pixel, at least one neighboring pixel, and at least one non-neighboring pixel in at least one reference block, thereby ensuring the effectiveness of the spatial features. When the spatial features are used to perform prediction processing on at least one image block, key image block features can be focused, thereby improving the prediction effect on at least one image block.
[0249] In the eighth approach, at least one reference block includes a statistical feature value of a reference area of at least one of at least one reference pixel, at least one neighboring pixel, and at least one non-neighboring pixel.
[0250] Optionally, the reference block may include at least one reference area, and the reference area may include at least one reference sub-block. The concept of the reference sub-block may refer to the above-mentioned method seven.
[0251] Optionally, each reference area in a reference block may correspond to at least one spatial feature.
[0252] Optionally, a spatial feature of the reference block may be determined or obtained based on at least one reference region in at least one reference block. Optionally, the spatial feature of the reference block may include at least one spatial feature determined or obtained based on at least one reference region in the reference block.
[0253] Optionally, a reference area including at least one reference pixel or a first reference intermediate element, at least one neighboring pixel and at least one non-neighboring pixel can be determined from the reference block, and the spatial domain features of at least one reference block can be determined or obtained based on the statistical characteristic values of the reference area, and prediction processing can be performed on at least one image block based on the spatial domain features of at least one reference block.
[0254] Optionally, for all or part of the reference pixels or first reference intermediate elements in at least one reference block, a reference area in the at least one reference block containing at least one of the reference pixel or the first reference intermediate element, at least one neighboring pixel of the reference pixel or the first reference intermediate element, and at least one non-neighboring pixel of the reference pixel or the first reference intermediate element can be determined, and the spatial domain features of the at least one reference block can be determined or obtained based on the statistical characteristic values of the reference area, and prediction processing can be performed on the at least one image block based on the spatial domain features of the at least one reference block.
[0255] Optionally, a reference area including at least one reference pixel or a first reference intermediate element, at least one neighboring pixel and at least one non-neighboring pixel can be determined from the reference block, and the spatial domain features of at least one reference block can be determined or obtained based on the weighted statistical feature values of the reference area, and prediction processing can be performed on at least one image block based on the spatial domain features of at least one reference block.
[0256] In this embodiment, the spatial domain features are determined or obtained based on the statistical characteristic values of a reference area containing at least one reference pixel, at least one neighboring pixel, and at least one non-neighboring pixel in at least one reference block, thereby ensuring the validity of the spatial domain features. When the spatial domain features are used to perform prediction processing on at least one image block, key image block features can be focused, thereby improving the prediction effect on at least one image block.
[0257] Third embodiment
[0258] Based on any of the above embodiments, a third embodiment is proposed.
[0259] In this embodiment, step S1 includes at least one of the following steps a1 to a3:
[0260] Step a1: determining or obtaining at least one spatial attention weight based on at least one spatial feature and at least one lookup table, and performing prediction processing on at least one image block based on the at least one spatial attention weight;
[0261] Optionally, at least one spatial attention weight may be determined or obtained based on at least one spatial feature of at least one reference block and at least one lookup table, and prediction processing may be performed on at least one image block based on the at least one spatial attention weight.
[0262] Optionally, the spatial attention weight can be a self-attention weight at the spatial level.
[0263] Optionally, spatial attention weights are weight coefficients that quantify the importance of different spatial locations (regions) in the feature map. By learning or calculating the weight map, the model will pay more attention to high-weight regions (enhancing features) and pay less attention to low-weight regions (suppressing irrelevant information), thereby focusing on the most discriminative areas in the image (such as the target subject, key textures, etc.).
[0264] Alternatively, lookup tables are a common method for accelerating computation in embedded systems. For a complex function or series of calculations, if the output values are stored in a lookup table, all that needs to be done later is retrieve the values without having to perform the calculations. Therefore, lookup tables are effective when computation time is longer than memory access time.
[0265] Optionally, the lookup table can be used to approximate the calculation of a neural network or a neural network module, such as a 3x3 convolutional layer. The neural network or the neural network module is pre-trained, and the lookup table is obtained based on the input-output mapping of the trained model.
[0266] Optionally, the lookup table may include at least one spatial feature, at least one spatial attention weight, and a corresponding relationship between the two. Optionally, the lookup table may also include at least one spatial feature, at least one intermediate element, and a corresponding relationship between the two. Optionally, the lookup table may also include at least one intermediate element, at least one spatial attention weight, and a corresponding relationship between the two.
[0267] Optionally, at least one spatial attention weight may be determined or obtained in at least one lookup table based on at least one spatial feature of at least one image block.
[0268] Optionally, after determining at least one spatial feature of at least one image block, the at least one spatial feature can be input into at least one lookup table for search. Since the lookup table includes the correspondence between at least one spatial feature and at least one spatial attention weight, the spatial attention weight that has a correspondence with the at least one spatial feature of the input can be searched in the at least one lookup table and output, and then the at least one image block is predicted based on the output at least one spatial attention weight.
[0269] Optionally, at least one spatial feature may be input into at least one lookup table for search to obtain at least one intermediate element, and at least one spatial attention weight may be determined or obtained based on the at least one intermediate element. For example, the at least one intermediate element may be input into at least one neural network and / or at least one activation function to output at least one spatial attention weight.
[0270] Optionally, at least one intermediate element can be determined or obtained based on at least one spatial feature, for example, at least one spatial feature is input into at least one neural network and / or at least one activation function, at least one intermediate element is output, and the intermediate element is input into at least one lookup table for search to obtain at least one spatial attention weight.
[0271] Optionally, the reference block may include reference block features of at least one channel. Optionally, a spatial attention weight corresponding to the reference block features of at least one channel may be determined or obtained based on the spatial features of the reference block and at least one lookup table, and prediction processing may be performed on the image block based on the spatial attention weight corresponding to the reference block features of the channel.
[0272] In this embodiment, at least one spatial attention weight is determined or obtained based on at least one spatial feature of at least one reference block and at least one lookup table, and prediction processing is performed on at least one image block based on the at least one spatial attention weight. The key image block features that need to be processed in the image block can be determined by the spatial attention weight, and the key image block features can be focused during the prediction processing, which can improve the prediction effect of the prediction processing on at least one image block.
[0273] Step a2: determining or obtaining at least one spatial attention weight based on at least one spatial feature and at least one neural network, and performing prediction processing on at least one image block based on the at least one spatial attention weight;
[0274] Optionally, at least one spatial attention weight can be determined or obtained based on at least one spatial feature of at least one reference block and at least one neural network, and prediction processing can be performed on at least one image block based on the at least one spatial attention weight.
[0275] Optionally, the spatial attention weight can be a self-attention weight at the spatial level.
[0276] Optionally, the neural network in this embodiment can be a neural network based on a fully connected layer, a neural network based on a convolutional layer, a neural network based on a Transformer, a neural network based on a hybrid convolutional layer, a fully connected layer and a Transformer, etc.
[0277] Optionally, at least one spatial domain feature may be input into at least one neural network, and at least one spatial attention weight may be output.
[0278] Optionally, at least one spatial feature may be input into at least one neural network, at least one intermediate element may be output, and at least one spatial attention weight may be determined or obtained based on the at least one intermediate element. For example, at least one intermediate element may be input into at least one lookup table or at least one activation function, and at least one spatial attention weight may be output.
[0279] Optionally, at least one intermediate element can be determined or obtained based on at least one spatial feature, for example, at least one spatial feature is input into at least one lookup table or at least one activation function, at least one intermediate element is output, the intermediate element is input into at least one neural network, and at least one spatial attention weight is output.
[0280] Optionally, at least one spatial domain feature may be input into at least one lookup table for search, an intermediate element may be output, and the intermediate element may be input into at least one neural network to output at least one spatial attention weight.
[0281] Optionally, the lookup table may include at least one spatial feature, at least one intermediate element, and a correspondence between the two. Since the lookup table includes the correspondence between the at least one spatial feature and the at least one intermediate element, the intermediate element that corresponds to the at least one input spatial feature may be searched in the at least one lookup table and output. The at least one intermediate element output is then input into at least one neural network to obtain at least one spatial attention weight, and prediction processing is performed on the at least one image block based on the at least one spatial attention weight.
[0282] Optionally, at least one spatial domain feature may be input into at least one neural network, an intermediate element may be output, and the intermediate element may be input into at least one lookup table for search, and at least one spatial attention weight may be output.
[0283] Optionally, the lookup table may include at least one intermediate element, at least one spatial attention weight, and a corresponding relationship between the two. Since the lookup table includes the corresponding relationship between the at least one intermediate element and the at least one spatial attention weight, after inputting the at least one spatial feature into the at least one neural network and outputting the intermediate element, the at least one intermediate element may be input into the at least one lookup table to determine at least one spatial attention weight that corresponds to the input at least one intermediate element, and prediction processing may be performed on the at least one image block based on the at least one spatial attention weight.
[0284] In this embodiment, at least one spatial attention weight is determined or obtained based on at least one spatial feature of at least one reference block and at least one neural network, and prediction processing is performed on at least one image block based on the at least one spatial attention weight. The key image block features that need to be processed in the reference block can be determined through spatial attention, and the key image block features can be focused during the prediction processing, which can improve the prediction effect of the prediction processing on at least one image block.
[0285] Step a3: Determine or obtain at least one spatial attention weight based on at least one spatial feature and at least one activation function, and perform prediction processing on at least one image block based on the at least one spatial attention weight.
[0286] Optionally, at least one spatial attention weight may be determined or obtained based on at least one spatial feature and at least one activation function of at least one reference block, and prediction processing may be performed on at least one image block based on the at least one spatial attention weight.
[0287] Optionally, the spatial attention weight can be a self-attention weight at the spatial level.
[0288] Optionally, an activation function, also known as an excitation function, often exists between the input layer and the output layer of a neural network. Its function is to add some nonlinear factors to the neural network.
[0289] Optionally, at least one spatial feature may be input into at least one activation function, and at least one spatial attention weight may be output.
[0290] Optionally, at least one spatial feature may be input into at least one activation function, at least one intermediate element may be output, and at least one spatial attention weight may be determined or obtained based on the at least one intermediate element. For example, at least one intermediate element may be input into at least one lookup table or at least one neural network, and at least one spatial attention weight may be output.
[0291] Optionally, at least one intermediate element can be determined or obtained based on at least one spatial feature, for example, at least one spatial feature is input into at least one lookup table or at least one neural network, at least one intermediate element is output, the intermediate element is input into at least one activation function, and at least one spatial attention weight is output.
[0292] Optionally, at least one spatial attention weight may be determined or obtained based on at least one spatial feature, at least one activation function, and at least one lookup table.
[0293] Optionally, at least one spatial domain feature may be input into at least one lookup table for search, an intermediate element may be output, and the intermediate element may be input into at least one activation function to output at least one spatial attention weight.
[0294] Optionally, the lookup table may include at least one spatial feature, at least one intermediate element, and a correspondence between the two. Since the lookup table includes the correspondence between the at least one spatial feature and the at least one intermediate element, the intermediate element that corresponds to the at least one input spatial feature can be searched in the at least one lookup table and output. The at least one intermediate element output is then input into at least one activation function to output at least one spatial attention weight.
[0295] Optionally, at least one spatial domain feature may be input into at least one activation function, an intermediate element may be output, and the intermediate element may be input into at least one lookup table for search, and at least one spatial attention weight may be output.
[0296] Optionally, the lookup table may include at least one intermediate element, at least one spatial attention weight, and a corresponding relationship between the two. Since the lookup table includes the corresponding relationship between the at least one intermediate element and the at least one spatial attention weight, after inputting the at least one spatial feature into the at least one activation function and outputting the intermediate element, the at least one intermediate element may be input into the at least one lookup table to search for and determine at least one spatial attention weight that has a corresponding relationship with the input at least one intermediate element.
[0297] Optionally, at least one spatial attention weight may be determined or obtained based on at least one spatial feature, at least one activation function, and at least one neural network.
[0298] Alternatively, at least one spatial feature may be input into at least one activation function, outputting an intermediate element, which is then input into at least one neural network to obtain at least one spatial attention weight. Alternatively, at least one spatial feature may be input into at least one neural network to obtain an intermediate element, which is then input into at least one activation function to obtain at least one spatial attention weight.
[0299] Optionally, at least one spatial attention weight may be determined or obtained based on at least one spatial feature, at least one activation function, at least one neural network, and at least one lookup table.
[0300] In this embodiment, at least one spatial attention weight is determined or obtained based on at least one spatial feature and at least one activation function of at least one reference block, and prediction processing is performed on at least one image block based on the at least one spatial attention weight. The key image block features that need to be processed in the reference block can be determined through spatial attention, and the key image block features can be focused during the prediction processing, which can improve the prediction effect of the prediction processing on at least one image block.
[0301] Fourth embodiment
[0302] Based on any of the above embodiments, a fourth embodiment is proposed.
[0303] In this embodiment, the image processing method further includes at least one of the following methods 9 to 12:
[0304] Mode nine: The spatial attention weights corresponding to at least two first intermediate elements in the intermediate element block corresponding to at least one reference block are the same and / or the spatial attention weights corresponding to the reference pixels are different;
[0305] Optionally, the intermediate element block may include at least one intermediate element, such as a first intermediate element. The first intermediate element may be an intermediate element. The intermediate element may be an element determined based on a reference pixel in a reference block. The intermediate element may be a reference pixel. The intermediate pixel may be a pixel obtained by filtering, transforming, downsampling, or downsampling the reference pixel.
[0306] Optionally, the intermediate element block may be an intermediate element block including at least one intermediate element obtained by pre-filtering, transforming, down-sampling, or down-sampling at least one reference pixel in at least one reference block.
[0307] Optionally, for at least two first intermediate elements in at least one intermediate element block, spatial attention weights corresponding to the at least two first intermediate elements can be determined or obtained, and the spatial attention weights corresponding to the at least two first intermediate elements can be different.
[0308] For at least two first intermediate elements in at least one intermediate element block, the spatial attention weights corresponding to the at least two first intermediate elements can be determined or obtained based on at least one spatial feature and at least one lookup table. The spatial attention weights corresponding to the at least two first intermediate elements may be different, and the at least one intermediate element block is predicted based on the spatial attention weights corresponding to the at least two first intermediate elements.
[0309] For example, if the size parameter of the intermediate element block is 3×3, the spatial attention weights corresponding to at least two of the nine first intermediate elements in the intermediate element block may be different. For example, the spatial attention weights corresponding to the nine first intermediate elements in the image block are all different.
[0310] Optionally, the spatial attention weights corresponding to the first intermediate elements of at least two rows or at least two columns in the intermediate element block may be different. For example, the spatial attention weight corresponding to the first intermediate element of the first row may be different from the spatial attention weight corresponding to the first intermediate element of the second row, or the spatial attention weights corresponding to the first intermediate elements of each row may be different.
[0311] Optionally, for at least two first intermediate elements in at least one intermediate element block, spatial attention weights corresponding to the at least two first intermediate elements can be determined or obtained, and the spatial attention weights corresponding to the at least two first intermediate elements can be the same.
[0312] For example, if the size parameter of the intermediate element block is 3×3, the spatial attention weights corresponding to at least two of the nine first intermediate elements in the intermediate element block may be the same. For example, the spatial attention weights corresponding to the nine first intermediate elements in the intermediate element block are all the same.
[0313] Optionally, the spatial attention weights corresponding to the first intermediate elements of at least two rows or at least two columns in the intermediate element block may be the same. For example, the spatial attention weight corresponding to the first intermediate element of the first row may be the same as the spatial attention weight corresponding to the first intermediate element of the second row, or the spatial attention weights corresponding to the first intermediate elements of each row may be the same.
[0314] Optionally, for at least three first intermediate elements in at least one intermediate element block, spatial attention weights corresponding to the at least three first intermediate elements can be determined or obtained, wherein the spatial attention weights corresponding to at least two first intermediate elements may be different, and the spatial attention weights corresponding to at least two first intermediate elements may be the same.
[0315] For example, if the size parameter of the intermediate element block is 3×3, the spatial attention weights corresponding to at least two of the nine first intermediate elements in the intermediate element block can be the same, and the spatial attention weights corresponding to at least two of the first intermediate elements can be different. For example, the spatial attention weights corresponding to four of the first intermediate elements in the intermediate element block are all the same, and the spatial attention weights corresponding to the other five first intermediate elements are different.
[0316] Optionally, the spatial attention weights corresponding to the first intermediate elements of at least two rows or at least two columns in the intermediate element block may be the same, and the spatial attention weights corresponding to the first intermediate elements of at least two rows or at least two columns may be different. For example, the spatial attention weight corresponding to the first intermediate element of the first row may be the same as the spatial attention weight corresponding to the first intermediate element of the second row, or the spatial attention weight corresponding to the first intermediate element of the second row may be different from the spatial attention weight corresponding to the first intermediate element of the third row.
[0317] Optionally, the spatial attention weights corresponding to at least two first intermediate elements in at least one intermediate element block can be determined or obtained based on at least one spatial feature of at least one intermediate element block and at least one lookup table, at least one activation function and at least one neural network. The spatial attention weights corresponding to the at least two first intermediate elements may be different, and prediction processing is performed on at least one image block based on the different spatial attention weights corresponding to the at least two first intermediate elements.
[0318] In this embodiment, for at least two first intermediate elements in at least one intermediate element block, spatial attention weights corresponding to the at least two first intermediate elements can be determined or obtained, and the spatial attention weights corresponding to the at least two first intermediate elements can be different and / or the same. At least one image block is predicted based on the at least one spatial attention weight. The key image block features that need to be processed in the intermediate element block can be determined by different and / or the same spatial attention weights, and thus the key image block features can be focused during the prediction processing, which can improve the prediction effect of the prediction processing on at least one image block.
[0319] Method 10: The spatial attention weight corresponding to at least one first intermediate element in the intermediate element block corresponding to at least one reference block is greater than the spatial attention weight corresponding to another reference pixel;
[0320] Optionally, for at least two first intermediate elements in at least one intermediate element block, spatial attention weights corresponding to the at least two first intermediate elements can be determined or obtained, wherein the spatial attention weight corresponding to at least one first intermediate element is greater than the spatial attention weight corresponding to the other first intermediate element.
[0321] Optionally, for at least two first intermediate elements in at least one intermediate element block, spatial attention weights corresponding to the at least two first intermediate elements can be determined or obtained based on at least one spatial feature and at least one lookup table, wherein the spatial attention weight corresponding to at least one first intermediate element is greater than the spatial attention weight corresponding to another first intermediate element, and prediction processing is performed on at least one image block based on the spatial attention weights corresponding to the at least two first intermediate elements.
[0322] For example, if the size parameter of the intermediate element block is 3×3, the spatial attention weight corresponding to at least one of the nine first intermediate elements in the intermediate element block is greater than the spatial attention weight corresponding to another first intermediate element. For example, the spatial attention weight corresponding to one first intermediate element in the intermediate element block is greater than the spatial attention weights corresponding to the other eight first intermediate elements.
[0323] Optionally, the spatial attention weight corresponding to the first intermediate element of at least one row or column in the intermediate element block may be greater than the spatial attention weight corresponding to the first intermediate element of another row or column. For example, the spatial attention weight corresponding to the first intermediate element of the first row is greater than the spatial attention weight corresponding to the first intermediate element of the second row.
[0324] Optionally, the spatial attention weights corresponding to at least two first intermediate elements in at least one intermediate element block can be determined or obtained based on at least one spatial feature and at least one lookup table, at least one activation function and at least one neural network. Optionally, the spatial attention weight corresponding to at least one first intermediate element is greater than the spatial attention weight corresponding to another first intermediate element. Based on the spatial attention weights corresponding to the at least two first intermediate elements, prediction processing is performed on at least one image block.
[0325] In this embodiment, for at least two first intermediate elements in at least one intermediate element block, spatial attention weights corresponding to the at least two first intermediate elements can be determined or obtained, and the spatial attention weight corresponding to at least one first intermediate element is greater than the spatial attention weight corresponding to another first intermediate element. At least one image block is predicted based on the at least one spatial attention weight. The key image block features that need to be processed in the intermediate element block can be determined by spatial attention weights of different sizes, and the key image block features can be focused during the prediction processing, which can improve the prediction effect of the prediction processing on at least one image block.
[0326] In an eleventh embodiment, a spatial attention weight corresponding to an intermediate element region containing at least one first intermediate element in an intermediate element block corresponding to at least one reference block is greater than a spatial attention weight corresponding to an intermediate element region containing at least one other first intermediate element;
[0327] Optionally, the intermediate element block includes at least one intermediate element region, wherein the intermediate element region includes at least one first intermediate element. Optionally, an intermediate element region may correspond to at least one spatial attention weight, or at least one intermediate element region may correspond to one spatial attention weight.
[0328] Optionally, for at least two intermediate element regions in at least one intermediate element block, spatial attention weights corresponding to the at least two intermediate element regions can be determined or obtained, wherein the spatial attention weight corresponding to the intermediate element region containing at least one first intermediate element is greater than the spatial attention weight corresponding to the intermediate element region containing at least one other first intermediate element.
[0329] Optionally, for at least two intermediate element regions in at least one intermediate element block, spatial attention weights corresponding to the at least two intermediate element regions can be determined or obtained based on at least one spatial domain feature and at least one lookup table, wherein the spatial attention weight corresponding to the intermediate element region containing at least one first intermediate element is greater than the spatial attention weight corresponding to the intermediate element region containing at least one other first intermediate element, and prediction processing is performed on at least one image block based on the spatial attention weights corresponding to the at least two intermediate element regions.
[0330] For example, the size parameter of the intermediate element block is 3×3, and the intermediate element block contains two intermediate element areas, one of which contains the first intermediate elements of the first and second rows, and the other intermediate element area contains the first intermediate element of the third row. The spatial attention weight corresponding to the intermediate element area containing the first intermediate elements of the first and second rows can be greater than the spatial attention weight corresponding to the intermediate element area containing the first intermediate element of the third row.
[0331] Optionally, the spatial attention weights corresponding to at least two intermediate element regions in at least one intermediate element block can be determined or obtained based on at least one spatial domain feature and at least one lookup table, at least one activation function and at least one neural network, wherein the spatial attention weight corresponding to the intermediate element region containing at least one first intermediate element is greater than the spatial attention weight corresponding to the intermediate element region containing at least one other first intermediate element, and prediction processing is performed on at least one image block based on the spatial attention weights corresponding to the at least two intermediate element regions.
[0332] In this embodiment, for at least two intermediate element regions in at least one intermediate element block, spatial attention weights corresponding to the at least two intermediate element regions can be determined or obtained, wherein the spatial attention weight corresponding to the intermediate element region containing at least one first intermediate element is greater than the spatial attention weight corresponding to the intermediate element region containing at least one other first intermediate element. At least one image block is predicted based on the at least one spatial attention weight, and the key image block features that need to be processed in the intermediate element block can be determined by using spatial attention weights of different sizes, thereby focusing on the key image block features during the prediction processing, thereby improving the prediction effect of the prediction processing on at least one intermediate element block.
[0333] Method 12: The spatial attention weight corresponding to the intermediate element sub-block containing at least one first intermediate element in the intermediate element block corresponding to at least one reference block is greater than the spatial attention weight corresponding to the intermediate element sub-block containing at least one other first intermediate element.
[0334] Optionally, the intermediate element block includes at least one intermediate element region, the intermediate element region includes at least one intermediate element sub-block, and the intermediate element sub-block includes at least one first intermediate element.
[0335] Optionally, at least one intermediate element sub-block may correspond to at least one spatial attention weight, or at least one intermediate element sub-block may correspond to a spatial attention weight.
[0336] Optionally, for at least two intermediate element sub-blocks in at least one intermediate element block, spatial attention weights corresponding to the at least two intermediate element sub-blocks can be determined or obtained, wherein the spatial attention weight corresponding to the intermediate element sub-block containing at least one first intermediate element is greater than the spatial attention weight corresponding to the intermediate element sub-block containing at least one other first intermediate element.
[0337] Optionally, for at least two intermediate element sub-blocks in at least one intermediate element block, spatial attention weights corresponding to the at least two intermediate element sub-blocks can be determined or obtained based on at least one spatial domain feature and at least one lookup table, the spatial attention weight corresponding to the intermediate element sub-block containing at least one first intermediate element is greater than the spatial attention weight corresponding to the intermediate element sub-block containing at least one other first intermediate element, and prediction processing is performed on at least one image block based on the spatial attention weights corresponding to the at least two intermediate element sub-blocks.
[0338] For example, the size parameter of the intermediate element block is 5×5. For a 3×3 intermediate element area contained in the intermediate element block, the intermediate element area contains two intermediate element sub-blocks, one of which contains the first intermediate elements of the first and second rows in the intermediate element area, and the other intermediate element sub-block contains the first intermediate element of the third row in the intermediate element area. The spatial attention weight corresponding to the intermediate element sub-block containing the first intermediate elements of the first and second rows can be greater than the spatial attention weight corresponding to the intermediate element sub-block containing the first intermediate element of the third row.
[0339] Optionally, the spatial attention weights corresponding to at least two intermediate element sub-blocks in at least one intermediate element block can be determined or obtained based on at least one spatial feature of at least one intermediate element block and at least one lookup table, at least one activation function and at least one neural network, wherein the spatial attention weight corresponding to the intermediate element sub-block containing at least one first intermediate element is greater than the spatial attention weight corresponding to the intermediate element sub-block containing at least one other first intermediate element, and prediction processing is performed on at least one image block based on the spatial attention weights corresponding to the at least two intermediate element sub-blocks.
[0340] In this embodiment, for at least two intermediate element sub-blocks in at least one intermediate element block, spatial attention weights corresponding to the at least two intermediate element sub-blocks can be determined or obtained, wherein the spatial attention weight corresponding to the intermediate element sub-block containing at least one first intermediate element is greater than the spatial attention weight corresponding to the intermediate element sub-block containing at least one other first intermediate element. At least one image block is predicted based on the at least one spatial attention weight, and key image block features that need to be processed in the intermediate element block can be determined by using spatial attention weights of different sizes, thereby focusing on the key image block features during the prediction processing, thereby improving the prediction effect of the prediction processing on at least one image block.
[0341] Fifth embodiment
[0342] Based on any of the above embodiments, a fifth embodiment is proposed.
[0343] In this embodiment, at least one intermediate element block is determined or obtained based on at least one neural network, a lookup table structure, at least one item in the lookup table, and at least one reference block;
[0344] Alternatively, the neural network may be a neural network module, such as a 3x3 convolutional layer.
[0345] Optionally, when performing prediction processing on at least one image block, at least one reference block may be input into a neural network for processing to determine or obtain at least one intermediate element block, and then subsequent prediction processing may be performed based on the at least one intermediate element block.
[0346] Optionally, a search index corresponding to at least one reference block may be determined, searched in at least one lookup table according to the index, at least one intermediate element block may be determined or obtained based on the search result, and subsequent prediction processing may be performed based on the at least one intermediate element block.
[0347] Optionally, at least one reference block may be processed according to at least one lookup table structure to determine or obtain at least one intermediate element block, and subsequent prediction processing may be performed based on the at least one intermediate element block.
[0348] Optionally, at least one reference block may be pre-processed, such as a relatively simple transformation, to determine or obtain at least one intermediate element block, and then subsequent prediction processing may be performed based on the at least one intermediate element block.
[0349] Optionally, the reference block includes at least one reference pixel.
[0350] Optionally, the first reference intermediate element may be an intermediate element, and the first intermediate element may be an intermediate element, and the stages of the two may be different. Optionally, the intermediate element may refer to the above description.
[0351] Optionally, the lookup table structure includes at least one of the following modes 13 to 29:
[0352] Method 13, at least one lookup table;
[0353] Method 14, at least one neural network module;
[0354] Optionally, the neural network module can be a neural network, such as a neural network that only includes a 3x3 convolutional layer, or a neural network module that includes at least one convolutional layer.
[0355] Optionally, the lookup table structure may include at least one lookup table, may include at least one neural network module, or may include at least one lookup table and at least one neural network module at the same time.
[0356] Optionally, the index corresponding to at least one reference pixel or the first reference intermediate element can be determined, and the index can be input into at least one lookup table in the lookup table structure for search, and the predicted block pixel or predicted intermediate element can be determined or obtained based on the search result; at least one reference pixel or the first reference intermediate element can be input into the neural network module, and the predicted pixel or predicted intermediate element can be output.
[0357] Optionally, the predicted pixel may be a pixel obtained by predicting the pixel to be predicted in the image block.
[0358] Optionally, an index corresponding to at least one reference pixel or a first reference intermediate element can be determined, and the index can be input into at least one lookup table in a lookup table structure for search, and the search result can be input into a neural network module for processing. Based on the output of the neural network module, a predicted pixel after prediction processing is determined or obtained, such as directly outputting the predicted pixel, or outputting a predicted intermediate element, and then inputting the predicted intermediate element into at least one lookup table structure (such as a lookup table or a neural network module) for processing, and finally obtaining the predicted pixel. The predicted intermediate element can be an intermediate output generated during the prediction process, such as a lookup table index, an intermediate result after downsampling, a pixel statistical eigenvalue, an autocorrelation matrix, such as an eigenvalue output by a neural network module, an attention weight, etc.
[0359] Optionally, at least one reference pixel or a first reference intermediate element can be input into the neural network module, at least one index can be determined based on the output result of the neural network module, and the at least one index can be input into at least one lookup table for search, and the predicted pixel can be determined or obtained based on the search result. For example, the search result is a predicted pixel or a eigenvalue or a prediction mode, and the predicted pixel is predicted based on the eigenvalue or the prediction mode. It can also be a predicted intermediate element, and processing is continued based on the predicted intermediate element in combination with at least one lookup table structure to finally obtain the predicted pixel.
[0360] In this embodiment, when the lookup table structure is at least one lookup table or at least one neural network module, at least one lookup table or at least one neural network module can be used to perform prediction processing on at least one pixel to be predicted, and then the lookup table or neural network module can be selected for prediction processing according to different scenarios to improve the prediction effect of the prediction processing, thereby supporting the improvement of the efficiency of video encoding and / or decoding.
[0361] Mode 15: at least two search branches, at least one search branch having the same input as another search branch;
[0362] Optionally, the lookup table structure may include at least two lookup branches, and each lookup branch may include at least one lookup table. The at least one lookup table included in each of the at least two lookup branches may be the same or different, which is not limited here.
[0363] Optionally, inputs of at least two search branches are the same, such as at least one of an input pixel value and an input pixel number.
[0364] Optionally, the inputs of at least one lookup table in at least one search branch and at least one lookup table in another search branch may be the same, for example, both come from the same predicted intermediate element block (such as the block to be predicted, or the reference image block determined or obtained based on the block to be predicted).
[0365] For example, the input of at least one lookup table in at least one search branch is two pixels, and the input of at least one lookup table in another search branch can also be two pixels. The two input pixels can be the pixels to be predicted, or they can be predicted intermediate elements determined or obtained based on the pixels to be predicted, etc.
[0366] In this embodiment, when the inputs of at least two search branches in the lookup table structure are the same, multiple prediction processes can be performed based on at least one pixel to be predicted in combination with at least two search branches with the same inputs in the lookup table structure, thereby improving the prediction effect.
[0367] Mode 16: at least two search branches, the input of at least one search branch is determined or obtained based on the output of another search branch;
[0368] Optionally, there are at least two search branches in the lookup table structure, and the inputs and outputs of the two search branches may have a certain correlation relationship.
[0369] Optionally, the input of at least one search branch can be determined or obtained based on the output of another search branch. For example, the output of at least one search branch can be used as the input of another search branch, the output range of at least one search branch can be used as the input range of another search branch, etc.
[0370] Optionally, the input of at least one lookup table in at least one search branch can be determined or obtained based on the output of at least one lookup table in another search branch. For example, the output of at least one lookup table in at least one search branch can be used as the input of at least one lookup table in another search branch, and the output range of at least one lookup table in at least one search branch can be used as the input range of at least one lookup table in another search branch, etc.
[0371] Optionally, the input of at least one search branch can be determined or obtained based on at least one reference pixel or the first reference intermediate element, for example, at least one reference pixel or the first reference intermediate element is used as the input of at least one search branch for prediction processing, or the index corresponding to at least one reference pixel or the first reference intermediate element is used as the input of at least one search branch for prediction processing, and then the output of the at least one search branch is used as the input of another search branch for prediction processing, until the predicted pixel is finally obtained.
[0372] In this embodiment, when the input of at least one search branch in the at least two search branches of the lookup table structure is determined or obtained based on the output of another search branch, at least one pixel to be predicted can be predicted in sequence based on the at least two search branches, thereby improving the prediction effect.
[0373] Mode 17: at least two search branches, the at least two search branches being located in at least one channel corresponding to at least one reference pixel or the first reference intermediate element;
[0374] Optionally, the channels in this embodiment may be applicable to a lookup table.
[0375] Optionally, a channel is a component of the feature map of an image block in the depth dimension, which is used to describe the feature representation of the number of features in a specific dimension. Each channel represents a certain feature (such as texture, edge, derivative distribution) extracted from an image block (such as a predicted block or a reconstructed block). For example, one channel may detect horizontal edges, and another channel may detect vertical edges.
[0376] Optionally, at least one reference pixel corresponds to multiple channels.
[0377] Optionally, at least one first predicted intermediate element corresponds to multiple channels.
[0378] Optionally, in the lookup table structure, there are at least two lookup branches, and the at least two lookup branches are located in the same channel.
[0379] Optionally, the at least two search branches may be used as a predictor-like device to perform prediction processing on at least one reference pixel or first reference intermediate element of the same channel.
[0380] Optionally, when processing pixel features (such as pixel value, pixel position, etc.) of at least one reference pixel or first reference intermediate element in the same channel, at least two search branches can be processed in parallel or serially, without limitation here.
[0381] Optionally, prediction processing may be performed on at least one reference pixel or first reference intermediate element in the same channel according to at least one lookup table in each of at least two search branches.
[0382] In this embodiment, by having at least two search branches in the lookup table structure, and at least two search branches being located in at least one channel corresponding to at least one reference pixel or the first reference intermediate element, the pixel features (such as pixel value or pixel position, etc.) of at least one reference pixel or the first reference intermediate element in the same channel can be predicted using at least two search branches, thereby embodying the realization of multiple prediction processing of pixel features of the same channel and improving the prediction effect.
[0383] Mode 18: at least two search branches, the search table search mode of at least two search branches is parallel search and / or serial search;
[0384] Optionally, the lookup table structure includes at least two search branches, and the search mode of the at least two search branches is parallel search and / or serial search.
[0385] Optionally, the input elements corresponding to the lookup tables in at least two search branches can be simultaneously determined based on at least one reference pixel or the first reference intermediate element, and input into their respective corresponding lookup tables for parallel search to obtain at least one predicted intermediate element, and then a subsequent prediction processing method is used to perform prediction processing to obtain a predicted pixel.
[0386] Optionally, the input element corresponding to the lookup table in at least one search branch can be determined or obtained based on at least one reference pixel or the first reference intermediate element, and input into the lookup table in the search branch for search. After the search is completed, a search operation of the lookup table in another search branch is performed, that is, a serial search is performed until the predicted pixel or predicted intermediate element is finally determined or obtained.
[0387] In this embodiment, by using parallel search and / or serial search as the search method of the lookup table in at least two search branches of the lookup table structure, it is possible to achieve that when at least two search branches of the lookup table structure are used to predict at least one pixel to be predicted, different search methods can be selected according to different scenarios to improve the prediction effect.
[0388] Mode 19: at least two lookup tables, wherein the input of at least one lookup table is determined or obtained according to the type, size parameter and / or input range of the other lookup table;
[0389] Optionally, there may be multiple types of lookup tables, the types of the multiple lookup tables may be the same, or the types of the multiple lookup tables may be different, such as a predicted pixel lookup table, a homography matrix lookup table, an eigenvalue lookup table, and the like.
[0390] Optionally, the size parameters of the lookup table may include the width, height, perimeter, and area of the lookup table.
[0391] Optionally, the input range of the lookup table may be set in advance, the input ranges of multiple lookup tables may be the same, or the input ranges of multiple lookup tables may be different.
[0392] Optionally, the lookup table structure includes at least two lookup tables, and the input of one lookup table can be determined according to the type of the other lookup table. For example, the types of two consecutive lookup tables are both predicted pixel lookup tables, and the input range of one predicted pixel lookup table is greater than or equal to the input range of the other predicted pixel lookup table.
[0393] Optionally, for at least two lookup tables in the lookup table structure, the input of one lookup table may be determined according to a size parameter of the other lookup table, for example, the input range of the lookup table with a larger size parameter is greater than or equal to the input range of the lookup table with a smaller size parameter.
[0394] Optionally, for at least two lookup tables in the lookup table structure, the input of one lookup table may be determined according to the input range of the other lookup table, for example, the inputs of the two lookup tables may be consistent.
[0395] In this embodiment, by determining or obtaining the input of at least one of the at least two lookup tables in the lookup table structure according to the type, size parameters and / or input range of the other lookup table, it is possible to perform multiple prediction processes on at least one pixel to be predicted using at least two lookup tables, thereby improving the prediction effect.
[0396] Mode 20: at least two lookup tables, wherein the at least two lookup tables are located in at least one channel corresponding to at least one reference pixel or the first reference intermediate element;
[0397] Optionally, at least one to-be-predicted pixel corresponds to multiple channels, and at least one first prediction intermediate element corresponds to multiple channels.
[0398] Optionally, the lookup table structure may include at least two lookup tables, and the at least two lookup tables may be respectively located in at least one channel of at least one pixel to be predicted, that is, there is at least one lookup table in each channel of at least one pixel to be predicted, or there may be at least two lookup tables in the same channel of at least one pixel to be predicted.
[0399] Optionally, at least two lookup tables can be respectively located in at least one channel of at least one first predicted intermediate element, that is, there is at least one lookup table in each channel of at least one first predicted intermediate element, or there can be at least two lookup tables in the same channel of at least one first predicted intermediate element.
[0400] In this embodiment, by having at least two lookup tables in the lookup table structure, at least two lookup tables are located in at least one channel corresponding to at least one reference pixel or the first reference intermediate element, it is possible to use at least two lookup tables to perform multiple prediction processes on at least one reference pixel or the first reference intermediate element of the same channel, thereby improving the prediction effect.
[0401] Mode 21: at least two lookup tables, the search mode of at least two lookup tables in the same channel is parallel search and / or serial search;
[0402] Optionally, in the lookup table structure, the search method of at least two lookup tables is parallel search. When searching at least two lookup tables in parallel, the inputs of the two lookup tables come from the same block (i.e., the reference block or the intermediate element block); by scanning the reference block or the intermediate element block once, the indexes of at least two lookup tables can be obtained at the same time, and then the at least two lookup tables can be searched in parallel based on the indexes of the at least two lookup tables.
[0403] Optionally, the intermediate element block may be a reference block containing predicted intermediate elements. The predicted intermediate elements may refer to the above description and will not be repeated here.
[0404] Optionally, in the lookup table structure, when searching at least two lookup tables in parallel, there is no dependency between the input and output of the two lookup tables, that is, the inputs of the at least two lookup tables can come from different blocks or from the same block; at least two lookup tables can be searched simultaneously by multiple tasks / processes / hardware.
[0405] Optionally, in the lookup table structure, at least two lookup tables are searched in serial mode. When searching at least two lookup tables serially, there is a dependency between the input and output of the two lookup tables. One of the lookup tables must be operated on first before the second lookup table can be operated on.
[0406] Optionally, the lookup table structure may include at least two lookup tables, and at least two lookup tables exist in the same channel corresponding to at least one pixel to be predicted, and the at least two lookup tables are searched in parallel and / or serially.
[0407] Optionally, at least two lookup tables exist in the same channel corresponding to at least one first predicted intermediate element, and the search mode of the at least two lookup tables is parallel search and / or serial search.
[0408] Optionally, when processing the predicted pixel according to the lookup table structure, at least two lookup tables in the same channel can be used to perform parallel and / or serial searches on the index corresponding to at least one reference pixel or the first reference intermediate element until the predicted pixel is finally obtained.
[0409] In this embodiment, by using parallel search and / or serial search as the search method for at least two lookup tables in the lookup table structure, at least two lookup tables in the same channel can be used. This allows different search methods to be selected according to different scenarios to improve the prediction effect when predicting at least one pixel to be predicted in the same channel using at least two search branches of the lookup table structure.
[0410] Mode 22: at least two lookup tables, and the lookup modes of the lookup tables in at least two channels are parallel search and / or serial search;
[0411] Optionally, the lookup table structure includes at least two lookup tables, and there are lookup tables in at least two channels corresponding to at least one pixel to be predicted, and the lookup tables in these two channels are searched in parallel and / or serially.
[0412] Optionally, there are lookup tables in at least two channels corresponding to at least one first predicted intermediate element, and the lookup tables in these two channels are searched in a parallel search and / or serial search manner.
[0413] Optionally, when processing the predicted pixel according to the lookup table structure, the lookup tables in at least two channels can be used to perform parallel searches on the indexes corresponding to the predicted pixels, or a serial search method can be used to call the lookup tables in at least two channels in sequence to perform searches until the predicted pixel is finally obtained.
[0414] Optionally, when processing the first predicted intermediate element according to the lookup table structure, the lookup tables in at least two channels can be used to perform parallel searches on the index corresponding to the first predicted intermediate element, or a serial search method can be used to call the lookup tables in at least two channels in sequence to perform searches until the predicted pixel is finally obtained.
[0415] In this embodiment, by using parallel search and / or serial search as the search method for the lookup tables in at least two channels in the at least two lookup tables of the lookup table structure, it is possible to achieve that when at least two search branches of the lookup table structure are used to perform prediction processing on at least one pixel to be predicted in different channels, different search methods can be selected according to different scenarios to improve the prediction effect.
[0416] Mode 23: at least one search branch and at least one neural network module, wherein the input of the at least one neural network module is the same as the input of the at least one search branch;
[0417] Optionally, the neural network module may include a neural network, such as at least one convolution layer, a 3x3 convolution layer, etc.
[0418] Optionally, an input of at least one search branch in the lookup table structure is the same as an input of at least one neural network module.
[0419] Optionally, at least one reference pixel or a first reference intermediate element can be simultaneously input into at least one search branch and at least one neural network module for parallel processing, and the output results of the parallel processing of at least one search branch and at least one neural network module can be fused or weighted, and then subsequent prediction processing can be performed until the predicted pixel is finally obtained.
[0420] In this embodiment, when the lookup table structure includes at least one search branch and at least one neural network module with the same input, it is possible to use at least one search branch of the lookup table and at least one neural network module to jointly perform prediction processing on at least one reference pixel or a first reference intermediate element, and combine the advantages of both the lookup table in the search branch and the neural network in the neural network module to perform prediction processing to improve the prediction effect.
[0421] Mode 24: at least one search branch and at least one neural network module, wherein an input of the at least one neural network module is determined or obtained based on an output of the at least one search branch;
[0422] Optionally, the input of at least one neural network module in the lookup table structure can be determined or obtained based on the output of at least one search branch, and the input of at least one neural network module can be determined or obtained based on the output of at least one lookup table.
[0423] Optionally, the index corresponding to at least one reference pixel or the first reference intermediate element can be input into the lookup table in at least one search branch for search to obtain the output of at least one search branch, and the output of at least one search branch can be used as the input of at least one neural network module (or the output of at least one search branch can be deformed, such as weighted processing, to obtain the input of at least one neural network module), and input into at least one neural network module for processing until all prediction processing processes are completed and the predicted pixel is obtained.
[0424] In this embodiment, when the lookup table structure includes at least one lookup branch and at least one neural network module, and the input of the at least one neural network module is determined or obtained based on the output of the at least one lookup branch, it is possible to use at least one lookup branch of the lookup table and at least one neural network module to jointly perform prediction processing on at least one reference pixel or the first reference intermediate element, and combine the advantages of both the lookup table in the lookup branch and the neural network in the neural network module to perform prediction processing to improve the prediction effect.
[0425] Mode 25: at least one lookup table and at least one neural network module, wherein the input of the at least one neural network module is determined or obtained according to the type, size parameter and / or input range of the at least one lookup table;
[0426] Optionally, the type, size parameters and / or input range of at least one lookup table can refer to the above method ten.
[0427] Optionally, the lookup table structure may include at least one lookup table and at least one neural network module. At least one reference pixel or a first reference intermediate element may be input into the at least one lookup table for search, and the input of at least one neural network module may be determined or obtained based on the search result, and input into at least one neural network module for processing until all prediction processing processes are completed and the predicted pixel is obtained.
[0428] Optionally, the input of at least one neural network module can be determined based on the type, size parameter and / or input range of at least one lookup table. For example, if the output ranges of two different types of lookup tables are different, their corresponding inputs to at least one neural network module will also be different; for example, if the output ranges of two lookup tables with different size parameters are different, their corresponding inputs to at least one neural network module will also be different.
[0429] In this embodiment, when the lookup table structure includes at least one lookup branch and at least one neural network module, and the input of the at least one neural network module is determined or obtained according to the type, size parameters and / or input range of the at least one lookup table, it is possible to use at least one lookup branch of the lookup table and at least one neural network module to jointly perform prediction processing on at least one reference pixel or the first reference intermediate element, and combine the advantages of both the lookup table in the lookup branch and the neural network in the neural network module to perform prediction processing, so as to improve the prediction effect.
[0430] Mode 26: at least two lookup tables, wherein the output ranges of the at least two lookup tables are the same and / or the output ranges of the lookup tables are different;
[0431] Optionally, output ranges of at least two lookup tables in the lookup table structure are the same.
[0432] Optionally, at least two lookup tables in the lookup table structure have different output ranges.
[0433] Optionally, there are at least three lookup tables in the lookup table structure, two lookup tables have the same output range, and two lookup tables have different output ranges.
[0434] For example, the lookup table structure includes lookup table 1, lookup table 2, and lookup table 3. Lookup table 1 and lookup table 2 have the same output range, and lookup table 3 has a different output range from lookup table 1 and from lookup table 2.
[0435] Optionally, at least two lookup tables with different output ranges may be used to predict at least one reference pixel or the first reference intermediate element; at least two lookup tables with the same output range may be used to predict at least one reference pixel or the first reference intermediate element.
[0436] In this embodiment, by including at least two lookup tables in the lookup table structure, and the output ranges of at least two lookup tables are the same and / or the output ranges of the lookup tables are different, it is possible to achieve more flexible prediction processing based on the output ranges of different lookup tables when predicting at least one reference pixel or the first reference intermediate element according to the lookup table structure, thereby improving the prediction effect.
[0437] Mode 27: at least two lookup tables, wherein the input ranges of the at least two lookup tables are the same and / or the input ranges of the lookup tables are different;
[0438] Optionally, at least two lookup tables in the lookup table structure have the same input range.
[0439] Optionally, at least two lookup tables in the lookup table structure have different input ranges.
[0440] Optionally, there are at least three lookup tables in the lookup table structure, two lookup tables have the same input range, and two lookup tables have different input ranges.
[0441] For example, the lookup table structure includes lookup table 1, lookup table 2, and lookup table 3. Lookup table 1 and lookup table 2 have the same input range, and lookup table 3 has a different input range from lookup table 1 and from lookup table 2.
[0442] Optionally, at least two lookup tables with different input ranges may be used to perform prediction processing on at least one reference pixel or the first reference intermediate element; and / or, at least two lookup tables with the same input range may be used to perform prediction processing on at least one reference pixel or the first reference intermediate element.
[0443] In this embodiment, by including at least two lookup tables in the lookup table structure, and the input ranges of at least two lookup tables are the same and / or the input ranges of the lookup tables are different, it is possible to achieve more flexible prediction processing based on the input ranges of different lookup tables when performing prediction processing on at least one reference pixel or the first reference intermediate element according to the lookup table structure, thereby improving the prediction effect.
[0444] Mode 28: at least one first residual module based on a lookup table, the first residual module comprising a first branch having at least one lookup table structure, an adder, and a second branch having a short-circuit structure;
[0445] Optionally, a residual structure may be provided in the lookup table structure, that is, at least one first residual module based on the lookup table may be provided, and the residual structure of the lookup table is embodied by the first residual module.
[0446] Optionally, a first residual module can be used to process at least one reference pixel or a first reference intermediate element, for example, the at least one reference pixel or the first reference intermediate element is processed respectively according to a first branch containing at least one lookup table structure and a second branch containing a short-circuit structure, and then an adder is used to perform weighted addition on the output results of the first branch and the second branch until all prediction processes are completed to obtain a predicted pixel.
[0447] In this embodiment, by including at least one first residual module based on a lookup table in the lookup table structure, the first residual module includes a first branch containing at least one lookup table structure, an adder, and a second branch containing a short-circuit structure, it can be achieved that when predicting at least one reference pixel or a first reference intermediate element based on the lookup table structure, the residual structure can be combined for prediction processing to reflect the advantages of the residual structure, thereby improving the prediction effect of the prediction processing.
[0448] Method 29: at least one second residual module based on a lookup table, the first input element of the second residual module passes through a first branch containing at least one lookup table structure to obtain a fourth output element, and the first input element of the third residual module passes through a second branch containing a short-circuit structure to obtain a fifth output element that is the same as the first input element, and the fourth output element and the fifth output element are weightedly added by an adder to obtain a sixth output element.
[0449] Optionally, a residual structure may be provided in the lookup table structure, such as at least one second residual module based on the lookup table, and / or at least one first residual module based on the lookup table.
[0450] Optionally, the first residual module may be the second residual module.
[0451] Optionally, the first input element (such as a pixel or an index mark, etc.) of the second residual module may be determined or obtained according to at least one pixel to be predicted.
[0452] Optionally, in the first branch of the second residual module, the first input element can be input into at least one lookup table structure for processing, such as inputting its corresponding index into the lookup table for search, and the obtained search result is used as the fourth output element of the first branch.
[0453] Optionally, in the second branch of the second residual module, the first input element may be input into the second branch containing the short-circuit result, and a fifth output element identical to the first input element may be output.
[0454] The fourth output element output by the first branch and the fifth output element output by the second branch are input into the adder for weighted addition to obtain a sixth output element.
[0455] Optionally, the sixth output element may be directly output as a predicted pixel, or the sixth output element may be used as a prediction intermediate element to continue subsequent prediction processing until all prediction processes are completed to obtain a predicted pixel.
[0456] In this embodiment, by including at least one second residual module based on the lookup table in the lookup table structure, the first input element of the second residual module passes through the first branch containing at least one lookup table structure to obtain a fourth output element, and the first input element of the third residual module passes through the second branch containing a short-circuit structure to obtain a fifth output element that is the same as the first input element, and the fourth output element and the fifth output element are weightedly added by an adder to obtain a sixth output element. This can achieve that when predicting at least one reference pixel or the first reference intermediate element according to the lookup table structure, the residual structure can be combined for predicting processing to reflect the advantages of the residual structure, thereby improving the prediction effect of the prediction processing.
[0457] Sixth embodiment
[0458] Based on any of the above embodiments, a sixth embodiment is proposed.
[0459] In this embodiment, the image processing method further includes: determining or obtaining spatial features of at least one reference block after image preprocessing.
[0460] Optionally, at least one spatial feature may be determined or obtained based on at least one reference block that has undergone image preprocessing, and prediction processing may be performed on at least one image block based on the at least one spatial feature.
[0461] Optionally, this embodiment may be combined with any method or step in the above embodiments.
[0462] In this embodiment, by determining or obtaining at least one spatial feature based on at least one reference block that has undergone image preprocessing, and performing prediction processing on at least one image block based on the at least one spatial feature, it is possible to comprehensively consider the spatial features of the at least one reference block when performing prediction processing on the at least one image block to determine the key image block features that need to be processed in the reference block, thereby focusing on the key image block features during prediction processing, and improving the prediction effect of the prediction processing on the at least one image block.
[0463] Optionally, the image preprocessing includes at least one of image sharpening, gradient calculation, transformation, and detection based on a detection operator.
[0464] Optionally, a spatial feature of at least one reference block after image sharpening may be determined or obtained, and prediction processing may be performed on at least one image block based on the at least one spatial feature.
[0465] Optionally, a spatial feature of at least one reference block after gradient calculation may be determined or obtained, and prediction processing may be performed on at least one image block based on the at least one spatial feature.
[0466] Optionally, a spatial feature of at least one transformed reference block may be determined or obtained, and prediction processing may be performed on at least one image block based on the at least one spatial feature.
[0467] Optionally, the transform can be a Fourier transform, a wavelet transform, or a deep learning based feature transform network.
[0468] Optionally, the transformation may also be a scale transformation, such as image scaling (bilinear interpolation, bicubic interpolation), multi-scale pyramid construction (such as Gaussian pyramid, Laplacian pyramid).
[0469] Optionally, the transformation may also be a linear transformation, a nonlinear transformation, or a projective transformation, such as a Hough transform or a perspective transform.
[0470] Optionally, a spatial feature of at least one reference block after texture detection using a detection operator may be determined or obtained, and prediction processing may be performed on at least one image block based on the at least one spatial feature.
[0471] Optionally, the detection operator includes at least one of a Sobel operator, a Scharr operator, a Canny operator, and a Laplacian operator.
[0472] Optionally, the detection operator may also be an operator based on a neural network.
[0473] The following is an example of image sharpening.
[0474] Optionally, for video compression tasks, since higher frequency information such as texture is lost more, 3x3 USM sharpening (unsharp mask sharpening) can be used to enhance the texture intensity of the image block as a priori input. The image after USM sharpening is recorded as right Extract spatial features at each position (i, j), such as Figure 7 shown.
[0475] Optionally, for position (i, j), extract a 5×5 window around it. In order to utilize the features of the entire window as much as possible, divide the window into inner and outer parts, that is, perform the operation of the Mean module in the figure, and then average the two areas to obtain the spatial domain features.
[0476] Optionally, calculate the feature mean of position (i, j) and its upper, lower, left, and right positions, that is, calculate the feature mean of position (i-1, j), position (i+1, j), position (i, j-1), and position (i, j+1), and record it as a inner , calculate the feature mean of the 8 outermost positions of the window, that is, calculate the feature mean of position (i-2,j-1), position (i-2,j+1), position (i-1,j-2), position (i+1,j-2), position (i+2,j-1), position (i+2,j+1), position (i-1,j+2), position (i+1,j+2), and record it as a exter ,(a inner , a exter ) is recorded as the spatial feature. According to the spatial feature of position (i, j), the spatial attention weight of position (i, j) can be further obtained:
[0477]
[0478] σ represents the sigmoid activation function; It can be a Conv1x2 layer or a lookup table. Conv1x2 indicates a 1x2 convolution kernel; style_feature is the spatial feature of the input, averaging the inner and outer features. In other words, to derive the spatial attention weight for position (i, j) based on the spatial features at position (i, j), a convolutional layer (Conv1x2) or a lookup table is used, followed by a sigmoid layer. Since the output of the sigmoid function ranges from 0 to 1, it can be used as an attention weight / mask. The lookup table can be obtained by storing the input and output of the convolutional layer (Conv1x2) in the lookup table.
[0479] Optionally, the result after spatial attention enhancement is:
[0480] f out =atten·f in +f in
[0481] f in It is the enhanced texture strength that serves as the prior input, and at the same time, the spatial attention weight is generated by spatial feature extraction, which further improves the network's ability to model spatial information and texture details.
[0482] In this embodiment, by performing prediction processing on at least one image block based on the spatial characteristics of at least one reference block after image preprocessing according to at least one of image sharpening, gradient calculation, transformation and detection operators for texture detection, the at least one image block is predicted based on the at least one spatial characteristic. This allows for comprehensive consideration of the spatial characteristics of the at least one reference block when predicting the at least one image block to determine key image block features that need to be processed in the reference block, thereby focusing on the key image block features during the prediction process, and improving the prediction effect of the prediction processing on the at least one image block.
[0483] Fifth embodiment
[0484] The present application also provides a processing device, referring to Figure 8 , the processing device includes:
[0485] The processing module A10 is configured to perform prediction processing on at least one image block according to the spatial domain features of at least one reference block.
[0486] Optionally, the spatial domain feature is determined or obtained according to at least one of the following:
[0487] Statistical feature values of a reference pixel and at least one neighboring pixel in at least one reference block;
[0488] Statistical feature values of at least two non-neighbor pixels in at least one reference block;
[0489] a statistical feature value of a reference pixel in at least one reference block and at least one of the above neighboring pixel, at least one left neighboring pixel, at least one right neighboring pixel, and at least one below neighboring pixel of the reference pixel;
[0490] a weighted statistical feature value of a reference pixel and pixels adjacent to the reference pixel in at least one reference block;
[0491] The statistical characteristic value of at least one pixel in the outermost layer of a window centered on the reference pixel, wherein the size parameter of the window is less than or equal to the size parameter of the reference block;
[0492] Statistical eigenvalues of eight pixels that are not adjacent to the reference pixel, and the eight pixels are not adjacent to each other;
[0493] a statistical characteristic value of a reference sub-block of at least one reference block including at least one of at least one reference pixel, at least one neighboring pixel, and at least one non-neighboring pixel;
[0494] The at least one reference block includes a statistical characteristic value of a reference area of at least one of at least one reference pixel, at least one neighboring pixel, and at least one non-neighboring pixel.
[0495] Optionally, the reference block is determined or obtained according to at least one of the following:
[0496] At least one of an upper adjacent pixel, an upper non-adjacent pixel, a left adjacent pixel, a left non-adjacent pixel, an upper left adjacent pixel, and an upper left non-adjacent pixel of the image block;
[0497] At least one of a neighbor block, a non-neighbor block, a cross-component block, a co-located block, a time domain block, and a default block corresponding to the image block;
[0498] at least one of the width, height, block size, and block area of the image block;
[0499] a candidate block determined or generated by a candidate motion vector or a candidate block vector of the image block;
[0500] If the first information of the image block satisfies the first condition, the reference block is the first reference block;
[0501] If the first information of the image block satisfies the second condition, the reference block is the second reference block.
[0502] Optionally, the processing module A10 is configured to perform at least one of the following:
[0503] Determining or obtaining at least one spatial attention weight based on at least one spatial feature and at least one lookup table, and performing prediction processing on at least one image block based on the at least one spatial attention weight;
[0504] Determining or obtaining at least one spatial attention weight based on at least one spatial feature and at least one neural network, and performing prediction processing on at least one image block based on the at least one spatial attention weight;
[0505] At least one spatial attention weight is determined or obtained based on at least one spatial feature and at least one activation function, and prediction processing is performed on at least one image block based on the at least one spatial attention weight.
[0506] Optionally, the processing module A10 is configured to perform at least one of the following:
[0507] The spatial attention weights corresponding to at least two first intermediate elements in the intermediate element block corresponding to at least one reference block are the same and / or the spatial attention weights corresponding to the reference pixels are different;
[0508] The spatial attention weight corresponding to at least one first intermediate element in the intermediate element block corresponding to at least one reference block is greater than the spatial attention weight corresponding to another reference pixel;
[0509] The spatial attention weight corresponding to the intermediate element region containing at least one first intermediate element in the intermediate element block corresponding to at least one reference block is greater than the spatial attention weight corresponding to the intermediate element region containing at least one other first intermediate element;
[0510] The spatial attention weight corresponding to the intermediate element sub-block containing at least one first intermediate element in the intermediate element block corresponding to at least one reference block is greater than the spatial attention weight corresponding to the intermediate element sub-block containing at least one other first intermediate element.
[0511] Optionally, the processing module A10 is configured to perform at least one of the following:
[0512] Determining or obtaining at least one intermediate element block based on at least one neural network, a lookup table structure, at least one item in the lookup table, and at least one reference block;
[0513] The lookup table structure includes at least one of the following:
[0514] at least one lookup table;
[0515] at least one neural network module;
[0516] At least two search branches, at least one search branch having the same input as another search branch;
[0517] at least two search branches, the input of at least one search branch being determined or obtained based on the output of another search branch;
[0518] At least two search branches, the at least two search branches being located in at least one channel corresponding to at least one reference pixel or a first reference intermediate element;
[0519] At least two search branches, the search tables of the at least two search branches are searched in a parallel search and / or serial search manner;
[0520] at least two lookup tables, wherein the input of at least one lookup table is determined or obtained according to the type, size parameter and / or input range of the other lookup table;
[0521] At least two lookup tables, the at least two lookup tables being located in at least one channel corresponding to at least one reference pixel or a first reference intermediate element;
[0522] At least two lookup tables, wherein the lookup mode of the at least two lookup tables in the same channel is parallel search and / or serial search;
[0523] At least two lookup tables, wherein the lookup tables in at least two channels are searched in parallel and / or serially;
[0524] at least one search branch and at least one neural network module, wherein an input of the at least one neural network module is the same as an input of the at least one search branch;
[0525] at least one search branch and at least one neural network module, wherein an input of the at least one neural network module is determined or obtained based on an output of the at least one search branch;
[0526] At least one lookup table and at least one neural network module, wherein an input of the at least one neural network module is determined or obtained according to a type, size parameter and / or input range of the at least one lookup table;
[0527] at least two lookup tables, wherein the output ranges of the at least two lookup tables are the same and / or the output ranges of the lookup tables are different;
[0528] at least two lookup tables, at least two of which have the same input range and / or different input ranges;
[0529] At least one first residual module based on a lookup table, the first residual module comprising a first branch including at least one lookup table structure, an adder, and a second branch including a short-circuit structure;
[0530] At least one second residual module based on a lookup table, a first input element of the second residual module passes through a first branch containing at least one lookup table structure to obtain a fourth output element, and a first input element of the third residual module passes through a second branch containing a short-circuit structure to obtain a fifth output element that is the same as the first input element, and the fourth output element and the fifth output element are weightedly added by an adder to obtain a sixth output element.
[0531] Optionally, performing prediction processing on at least one image block according to at least one spatial attention weight includes at least one of the following:
[0532] Determine or obtain at least one predicted block or predicted intermediate element according to a result of performing spatial attention enhancement on a first intermediate element in at least one intermediate element block according to at least one spatial attention weight;
[0533] Determine or obtain at least one predicted block or predicted intermediate element based on a result of performing spatial attention enhancement on at least one intermediate element region in at least one intermediate element block according to at least one spatial attention weight;
[0534] At least one predicted block or predicted intermediate element is determined or obtained based on the result of performing spatial attention enhancement on at least one intermediate element sub-block in at least one intermediate element block.
[0535] Optionally, the processing module A10 is configured to determine or obtain spatial features of at least one image block after image preprocessing.
[0536] Optionally, the image preprocessing includes at least one of the following:
[0537] Image sharpening;
[0538] Gradient calculation;
[0539] Transformation;
[0540] Detection is performed based on the detection operator.
[0541] The processing device provided in the embodiment of the present application has similar implementation principles and beneficial effects to the technical solutions shown in the above-mentioned corresponding method embodiments, and will not be described in detail here.
[0542] An embodiment of the present application further provides a processing device, including a memory and a processor. The memory stores an image processing program, and when the image processing program is executed by the processor, the steps of the image processing method in any of the above embodiments are implemented.
[0543] An embodiment of the present application further provides a storage medium on which an image processing program is stored. When the image processing program is executed by a processor, the steps of the image processing method in any of the above embodiments are implemented.
[0544] In the embodiments of the processing device and storage medium provided in this application, all technical features of any of the above-mentioned image processing method embodiments may be included. The expanded and explained contents of the specification are basically the same as those of the embodiments of the above-mentioned methods and will not be repeated here.
[0545] An embodiment of the present application further provides a computer program product, which includes computer program code. When the computer program code runs on a computer, the computer executes the methods in the various possible implementation modes described above.
[0546] An embodiment of the present application also provides a chip, including a memory and a processor, wherein the memory is used to store computer programs, and the processor is used to call and run the computer programs from the memory, so that a device equipped with the chip executes the methods in the various possible implementation modes as described above.
[0547] It is understood that the above scenarios are merely examples and do not limit the application scenarios of the technical solutions provided in the embodiments of this application. The technical solutions of this application can also be applied to other scenarios. For example, those skilled in the art will appreciate that with the evolution of system architecture and the emergence of new business scenarios, the technical solutions provided in the embodiments of this application will also be applicable to similar technical problems.
[0548] The serial numbers of the above embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.
[0549] The steps in the method of the embodiment of the present application can be adjusted in order, combined and deleted according to actual needs.
[0550] The units in the device of the embodiment of the present application can be merged, divided and deleted according to actual needs.
[0551] In this application, the same or similar terminology, technical solutions and / or application scenario descriptions are generally only described in detail the first time they appear. When they appear again later, they are generally not repeated for the sake of brevity. When understanding the technical solutions and other contents of this application, for the same or similar terminology, technical solutions and / or application scenario descriptions that are not described in detail later, you can refer to the previous relevant detailed descriptions.
[0552] In this application, the description of each embodiment has its own focus. For parts that are not described or recorded in detail in a certain embodiment, please refer to the relevant description of other embodiments.
[0553] The various technical features of the technical solution of this application can be combined arbitrarily. In order to make the description concise, not all possible combinations of the various technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.
[0554] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus the necessary general hardware platform, and of course can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) as mentioned above, and includes a number of instructions for enabling a terminal device (which can be a mobile phone, computer, server, controlled terminal, or network device, etc.) to execute the method of each embodiment of the present application.
[0555] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When software is used for implementation, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function according to the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a storage medium or transmitted from one storage medium to another storage medium. For example, the computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) means. The storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more available media integrations. The available medium can be a magnetic medium (e.g., a floppy disk, a storage disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state storage disk Solid State Disk (SSD)).
[0556] The above are only preferred embodiments of the present application and do not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made using the contents of the present application specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present application.
Claims
1. An image processing method, characterized in that: Including steps: S1, performing prediction processing on at least one image block according to spatial domain features of at least one reference block.
2. The image processing method according to claim 1, wherein: The airspace characteristics are determined or derived based on at least one of the following: Statistical feature values of a reference pixel and at least one neighboring pixel in at least one reference block; Statistical feature values of at least two non-neighbor pixels in at least one reference block; a statistical feature value of a reference pixel in at least one reference block and at least one of the above neighboring pixel, at least one left neighboring pixel, at least one right neighboring pixel, and at least one below neighboring pixel of the reference pixel; a weighted statistical feature value of a reference pixel and pixels adjacent to the reference pixel in at least one reference block; The statistical characteristic value of at least one pixel in the outermost layer of a window centered on the reference pixel, wherein the size parameter of the window is less than or equal to the size parameter of the reference block; Statistical eigenvalues of eight pixels that are not adjacent to the reference pixel, and the eight pixels are not adjacent to each other; a statistical characteristic value of a reference sub-block of at least one reference block including at least one of at least one reference pixel, at least one neighboring pixel, and at least one non-neighboring pixel; The at least one reference block includes a statistical characteristic value of a reference area of at least one of at least one reference pixel, at least one neighboring pixel, and at least one non-neighboring pixel.
3. The image processing method according to claim 1, wherein: The reference block is determined or obtained according to at least one of the following: At least one of an upper adjacent pixel, an upper non-adjacent pixel, a left adjacent pixel, a left non-adjacent pixel, an upper left adjacent pixel, and an upper left non-adjacent pixel of the image block; At least one of a neighbor block, a non-neighbor block, a cross-component block, a co-located block, a time domain block, and a default block corresponding to the image block; at least one of the width, height, block size, and block area of the image block; a candidate block determined or generated by a candidate motion vector or a candidate block vector of the image block; If the first information of the image block satisfies the first condition, the reference block is the first reference block; If the first information of the image block satisfies the second condition, the reference block is the second reference block.
4. The image processing method according to any one of claims 1 to 3, wherein: Step S1 includes at least one of the following: Determining or obtaining at least one spatial attention weight based on at least one spatial feature and at least one lookup table, and performing prediction processing on at least one image block based on the at least one spatial attention weight; Determining or obtaining at least one spatial attention weight based on at least one spatial feature and at least one neural network, and performing prediction processing on at least one image block based on the at least one spatial attention weight; At least one spatial attention weight is determined or obtained based on at least one spatial feature and at least one activation function, and prediction processing is performed on at least one image block based on the at least one spatial attention weight.
5. The image processing method according to claim 4, wherein: Also include at least one of the following: The spatial attention weights corresponding to at least two first intermediate elements in the intermediate element block corresponding to at least one reference block are the same and / or the spatial attention weights corresponding to the reference pixels are different; The spatial attention weight corresponding to at least one first intermediate element in the intermediate element block corresponding to at least one reference block is greater than the spatial attention weight corresponding to another reference pixel; The spatial attention weight corresponding to the intermediate element region containing at least one first intermediate element in the intermediate element block corresponding to at least one reference block is greater than the spatial attention weight corresponding to the intermediate element region containing at least one other first intermediate element; The spatial attention weight corresponding to the intermediate element sub-block containing at least one first intermediate element in the intermediate element block corresponding to at least one reference block is greater than the spatial attention weight corresponding to the intermediate element sub-block containing at least one other first intermediate element.
6. The image processing method according to claim 5, wherein: Also include at least one of the following: Determining or obtaining at least one intermediate element block based on at least one neural network, a lookup table structure, at least one item in the lookup table, and at least one reference block; The lookup table structure includes at least one of the following: at least one lookup table; at least one neural network module; At least two search branches, at least one search branch having the same input as another search branch; at least two search branches, the input of at least one search branch being determined or obtained based on the output of another search branch; At least two search branches, the at least two search branches being located in at least one channel corresponding to at least one reference pixel or a first reference intermediate element; At least two search branches, the search tables of the at least two search branches are searched in a parallel search and / or serial search manner; at least two lookup tables, wherein the input of at least one lookup table is determined or obtained according to the type, size parameter and / or input range of the other lookup table; At least two lookup tables, the at least two lookup tables being located in at least one channel corresponding to at least one reference pixel or a first reference intermediate element; At least two lookup tables, wherein the lookup mode of the at least two lookup tables in the same channel is parallel search and / or serial search; At least two lookup tables, wherein the lookup tables in at least two channels are searched in parallel and / or serially; at least one search branch and at least one neural network module, wherein an input of the at least one neural network module is the same as an input of the at least one search branch; at least one search branch and at least one neural network module, wherein an input of the at least one neural network module is determined or obtained based on an output of the at least one search branch; At least one lookup table and at least one neural network module, wherein an input of the at least one neural network module is determined or obtained according to a type, size parameter and / or input range of the at least one lookup table; at least two lookup tables, wherein the output ranges of the at least two lookup tables are the same and / or the output ranges of the lookup tables are different; at least two lookup tables, at least two of which have the same input range and / or different input ranges; At least one first residual module based on a lookup table, the first residual module comprising a first branch including at least one lookup table structure, an adder, and a second branch including a short-circuit structure; At least one second residual module based on a lookup table, a first input element of the second residual module passes through a first branch containing at least one lookup table structure to obtain a fourth output element, and a first input element of the third residual module passes through a second branch containing a short-circuit structure to obtain a fifth output element that is the same as the first input element, and the fourth output element and the fifth output element are weightedly added by an adder to obtain a sixth output element.
7. The image processing method according to claim 5, wherein: Performing prediction processing on at least one image block according to at least one spatial attention weight includes at least one of the following: Determine or obtain at least one predicted block or predicted intermediate element according to a result of performing spatial attention enhancement on a first intermediate element in at least one intermediate element block according to at least one spatial attention weight; Determine or obtain at least one predicted block or predicted intermediate element based on a result of performing spatial attention enhancement on at least one intermediate element region in at least one intermediate element block according to at least one spatial attention weight; At least one predicted block or predicted intermediate element is determined or obtained based on the result of performing spatial attention enhancement on at least one intermediate element sub-block in at least one intermediate element block.
8. The image processing method according to any one of claims 1 to 3, wherein: Also includes: Determine or obtain the spatial domain features of at least one reference block after image preprocessing.
9. The image processing method according to claim 8, wherein: Image preprocessing includes at least one of the following: Image sharpening; Gradient calculation; Transformation; Detection is performed based on the detection operator.
10. A processing device, characterized in that include: A memory and a processor, wherein an image processing program is stored in the memory, and when the image processing program is executed by the processor, the steps of the image processing method according to any one of claims 1 to 9 are implemented.
11. A storage medium, characterized in that: The storage medium stores a computer program, which, when executed by a processor, implements the steps of the image processing method according to any one of claims 1 to 9.
Citation Information
Patent Citations
Image component prediction method, encoder, decoder and storage medium
CN113676732A
Coding method, decoding method and related device
CN114079791A
Image processing method, device, equipment, system and storage medium
CN118901239A
Method and system for generating enhanced images
US20130051668A1
Improving the angle discretization in decoder side intra mode derivation
US20240380885A1