Image processing methods, processing devices and storage media
By using the spatial features of the reference block and spatial attention weights, the problem of not being able to focus on key image block features in neural network prediction processing is solved, and more efficient image block prediction results are achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SHENZHEN TRANSSION HLDG CO LTD
- Filing Date
- 2025-06-19
- Publication Date
- 2026-07-31
AI Technical Summary
In image processing, neural networks often fail to focus on key image patch features during prediction, resulting in poor prediction performance.
Image patches are predicted based on the spatial features of reference patches. Key image patch features are determined using spatial features and spatial attention weights, and prediction is performed using lookup tables and neural networks.
It improves the accuracy and efficiency of image patch prediction processing, focuses on key features, and enhances prediction results.
Smart Images

Figure CN120583237B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technology, specifically to an image processing method, processing device, and storage medium. Background Technology
[0002] Existing high-efficiency video coding frameworks, such as Neural Network Based Video Coding (NNVC) and / or Enhanced Compression Model (ECM), propose a video frame coding technique to improve coding performance without significantly increasing computational complexity.
[0003] In the process of conceiving and implementing this application, the inventors discovered at least the following problems:
[0004] In the prediction processing stage of the encoding and decoding process, neural networks can be used for prediction processing, such as NNIntraPrediction and NNInterPrediction. However, due to the large number of image patch features, the neural network needs to process multiple image patch features and cannot focus on key image patch features, resulting in poor prediction performance, which needs to be improved.
[0005] The preceding description is intended to provide general background information and does not necessarily constitute prior art. Summary of the Invention
[0006] To address the aforementioned technical problems, this application provides an image processing method, processing device, and storage medium, aiming to solve the technical problem of how to improve the prediction effect of prediction processing.
[0007] This application provides an image processing method, applicable to a processing device, comprising the following steps:
[0008] S1, perform prediction processing on at least one image block based on the spatial characteristics of at least one reference block.
[0009] Optionally, the spatial characteristics are determined or obtained based on at least one of the following:
[0010] Statistical feature values of a reference pixel and at least one neighboring pixel in at least one reference block;
[0011] Statistical feature values of at least two non-neighboring pixels in at least one reference block;
[0012] The statistical feature value of at least one of the following: a reference pixel in at least one reference block, at least one upper adjacent pixel, at least one left adjacent pixel, at least one right adjacent pixel, and at least one lower adjacent pixel of the reference pixel;
[0013] The reference pixel in at least one reference block, and the weighted statistical feature value of the pixels adjacent to the reference pixel;
[0014] The statistical feature value of at least one pixel of the outermost layer of the window centered on the reference pixel, where the window size parameter is less than or equal to the reference block size parameter;
[0015] The statistical feature values of eight pixels that are not adjacent to the reference pixel, and the eight pixels are not adjacent to each other;
[0016] The statistical feature value of a reference sub-block that contains at least one of the following: at least one reference pixel, at least one neighboring pixel, and at least one non-neighboring pixel;
[0017] At least one reference block contains statistical feature values of a reference region that include at least one reference pixel, at least one neighboring pixel, and at least one non-neighboring pixel.
[0018] Optionally, the reference block is determined or obtained based on at least one of the following:
[0019] At least one of the following: the top adjacent pixel, the top non-adjacent pixel, the left adjacent pixel, the left non-adjacent pixel, the top left adjacent pixel, and the top left non-adjacent pixel;
[0020] The image patch is corresponding to at least one of the following: neighboring block, non-neighboring block, cross-component block, co-position block, temporal block, and default block;
[0021] At least one of the following: width, height, block size, and block area of the image block;
[0022] Candidate motion vectors or candidate block vectors of image patches determine or generate candidate blocks;
[0023] If the first information of an image block satisfies the first condition, then the reference block is the first reference block;
[0024] If the first information of an image block satisfies the second condition, then the reference block is the second reference block.
[0025] Optionally, step S1 includes at least one of the following:
[0026] Based on at least one spatial feature and at least one lookup table, determine or obtain at least one spatial attention weight, and perform prediction processing on at least one image patch based on the at least one spatial attention weight.
[0027] Based on at least one spatial feature and at least one neural network, at least one spatial attention weight is determined or obtained, and at least one image patch is predicted based on the at least one spatial attention weight.
[0028] Based on at least one spatial feature and at least one activation function, at least one spatial attention weight is determined or obtained, and at least one image patch is predicted based on the at least one spatial attention weight.
[0029] Optionally, the image processing method further includes at least one of the following:
[0030] At least two first intermediate elements in the intermediate element block corresponding to at least one reference block have the same spatial attention weight and / or the spatial attention weights corresponding to the reference pixels are different.
[0031] The spatial attention weight corresponding to at least one first intermediate element in the intermediate element block corresponding to at least one reference block is greater than the spatial attention weight corresponding to another reference pixel;
[0032] The spatial attention weight corresponding to the intermediate element region of at least one reference block that contains at least one first intermediate element is greater than the spatial attention weight corresponding to the intermediate element region that contains at least one other first intermediate element.
[0033] The spatial attention weight corresponding to the intermediate element sub-block that contains at least one first intermediate element in the intermediate element block corresponding to at least one reference block is greater than the spatial attention weight corresponding to the intermediate element sub-block that contains at least one other first intermediate element.
[0034] Optionally, the image processing method further includes at least one of the following:
[0035] Based on at least one neural network, a lookup table structure, at least one item in the lookup table, and at least one reference block, determine or obtain at least one intermediate element block;
[0036] The lookup table structure includes at least one of the following:
[0037] At least one lookup table;
[0038] At least one neural network module;
[0039] There are at least two search branches, and at least one search branch has the same input as the other search branch;
[0040] There are at least two search branches, and the input of at least one search branch is determined or obtained based on the output of the other search branch;
[0041] At least two lookup branches, and at least two lookup branches are located in at least one channel corresponding to at least one reference pixel or the first reference intermediate element;
[0042] There are at least two search branches, and the search method for the search table with at least two search branches is parallel search and / or serial search;
[0043] At least two lookup tables, and the input of at least one lookup table is determined or obtained based on the type, size parameters and / or input range of the other lookup table;
[0044] At least two lookup tables, and at least two lookup tables are located in at least one channel corresponding to at least one reference pixel or the first reference intermediate element;
[0045] At least two lookup tables, and the lookup methods for at least two lookup tables within the same channel are parallel lookup and / or serial lookup;
[0046] At least two lookup tables, and the lookup tables in at least two channels are searched in parallel and / or serial manner;
[0047] At least one search branch and at least one neural network module, wherein the input of the at least one neural network module is the same as the input of the at least one search branch;
[0048] At least one search branch and at least one neural network module, wherein the input of the at least one neural network module is determined or obtained based on the output of the at least one search branch;
[0049] At least one lookup table and at least one neural network module, wherein the input of the at least one neural network module is determined or obtained based on the type, size parameters and / or input range of the at least one lookup table;
[0050] At least two lookup tables, at least two lookup tables have the same output range and / or the lookup tables have different output ranges;
[0051] At least two lookup tables, at least two lookup tables have the same input range and / or the lookup tables have different input ranges;
[0052] At least one first residual module based on a lookup table, the first residual module including a first branch containing at least one lookup table structure, an adder, and a second branch containing a short-circuit structure;
[0053] At least one second residual module based on a lookup table, the first input element of the second residual module is passed through a first branch containing at least one lookup table structure to obtain a fourth output element, and the first input element of the third residual module is passed through a second branch containing a short-circuit structure to obtain a fifth output element that is the same as the first input element. The fourth output element and the fifth output element are weighted and added by an adder to obtain a sixth output element.
[0054] Optionally, prediction processing is performed on at least one image patch based on at least one spatial attention weight, including at least one of the following:
[0055] Based on the spatial attention enhancement of the first intermediate element in at least one intermediate element block according to at least one spatial attention weight, at least one prediction block or prediction intermediate element is determined or obtained.
[0056] Based on at least one spatial attention weight, spatial attention enhancement is performed on at least one intermediate element region in at least one intermediate element block to determine or obtain at least one prediction block or prediction intermediate element.
[0057] Based on at least one spatial attention weight, the spatial attention enhancement is performed on at least one intermediate element sub-block in at least one intermediate element block to determine or obtain at least one prediction block or prediction intermediate element.
[0058] Optionally, the image processing method further includes: determining or obtaining the spatial features of at least one reference block after image preprocessing.
[0059] Optionally, image preprocessing includes at least one of the following:
[0060] Image sharpening;
[0061] Gradient calculation;
[0062] Transformation;
[0063] Detection is performed based on the detection operator.
[0064] This application also provides a processing apparatus, including:
[0065] The processing module is used to perform prediction processing on at least one image block based on the spatial features of at least one reference block.
[0066] This application also provides a processing device, including: a memory and a processor, wherein the memory stores an image processing program, and when the image processing program is executed by the processor, it implements the steps of any of the image processing methods described above.
[0067] This application also provides a storage medium storing a computer program that, when executed by a processor, implements the steps of any of the image processing methods described above.
[0068] As described above, the image processing method of this application can be applied to a processing device, including: performing predictive processing on at least one image block based on the spatial characteristics of at least one reference block. Through the technical solution of this application, when performing predictive processing on at least one image block, the spatial characteristics of at least one reference block are comprehensively considered to determine the key image block features that need to be processed within the reference block. This allows the predictive processing to focus on the key image block features, thereby improving the prediction effect when using at least one reference block for predictive processing. Attached Figure Description
[0069] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application. To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, those skilled in the art can obtain other drawings based on these drawings without any creative effort.
[0070] Figure 1 A schematic diagram of the hardware structure of a mobile terminal to implement the various embodiments of this application;
[0071] Figure 2 A communication network system architecture diagram provided for an embodiment of this application;
[0072] Figure 3 A schematic diagram of the hardware structure of a controller 140 provided in this application;
[0073] Figure 4 A schematic diagram of the hardware structure of a network node 150 provided in this application;
[0074] Figure 5 This is a flowchart illustrating the image processing method according to the first embodiment;
[0075] Figure 6 This is a flowchart illustrating the encoding and decoding process in image processing methods.
[0076] Figure 7 This is a schematic diagram of spatial feature extraction of an image patch in an image processing method;
[0077] Figure 8 This is a schematic diagram of the processing module of the processing device.
[0078] The realization of the objectives, functional features, and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. The accompanying drawings have illustrated specific embodiments of this application, which will be described in more detail below. These drawings and textual descriptions are not intended to limit the scope of the concept in any way, but rather to illustrate the concepts of this application to those skilled in the art through reference to specific embodiments. Detailed Implementation
[0079] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.
[0080] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, components, features, and elements with the same names in different embodiments of this application may have the same meaning or different meanings, the specific meaning of which must be determined by its interpretation in that specific embodiment or further in conjunction with the context of that specific embodiment.
[0081] It should be understood that although the terms first, second, third, etc., may be used herein to describe various information, such information should not be limited to these terms. These terms are used only to distinguish information of the same type from one another. For example, without departing from the scope of this document, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the word “if” as used herein may be interpreted as “when…” or “in response to determination”. Furthermore, as used herein, the singular forms “a,” “an,” and “the” are intended to also include the plural forms unless the context indicates otherwise. It should be further understood that the terms “comprising,” “including,” indicate the presence of the stated feature, step, operation, element, component, item, kind, and / or group, but do not exclude the presence, occurrence, or addition of one or more other features, steps, operations, elements, components, items, kinds, and / or groups. The terms “or,” “and / or,” “including at least one of the following,” etc., as used in this application may be interpreted as inclusive, or mean any one or any combination thereof. For example, "including at least one of the following: A, B, C" means "any one of the following: A; B; C; A and B; A and C; B and C; A and B and C." Similarly, "A, B, or C" or "A, B, and / or C" means "any one of the following: A; B; C; A and B; A and C; B and C; A and B and C." Exceptions to this definition only occur when the combination of elements, functions, steps, or operations is inherently mutually exclusive in some way.
[0082] It should be understood that although the steps in the flowcharts of this application's embodiments are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some of the steps in the figures may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be performed alternately or in turn with other steps or at least a portion of the sub-steps or stages of other steps.
[0083] Depending on the context, the words “if” or “suppose” as used here can be interpreted as “when” or “in response to determination” or “in response to detection.” Similarly, depending on the context, the phrases “if determination” or “if detection (of the stated condition or event)” can be interpreted as “when determination” or “in response to determination” or “when detection (of the stated condition or event)” or “in response to detection (of the stated condition or event).”
[0084] It should be noted that step designations such as S1 are used in this paper for the purpose of more clearly and concisely describing the corresponding content, and do not constitute a substantial restriction on the order.
[0085] It should be understood that the specific embodiments described herein are merely illustrative of this application and are not intended to limit this application.
[0086] In the following description, the use of suffixes such as "module," "part," or "unit" to denote elements is solely for the purpose of illustrative purposes and has no specific meaning in itself. Therefore, "module," "part," or "unit" may be used interchangeably.
[0087] The processing device in this application can be a smart terminal or a server, and the smart terminal can be implemented in various forms. For example, the smart terminal can include smart terminals such as mobile phones, tablets, laptops, handheld computers, personal digital assistants (PDAs), portable media players (PMPs), navigation devices, wearable devices, smart bracelets, pedometers, etc., as well as fixed terminals such as digital TVs and desktop computers.
[0088] The following description will use a mobile terminal as an example. Those skilled in the art will understand that, apart from elements specifically designed for mobile purposes, the construction according to the embodiments of this application can also be applied to fixed-type terminals.
[0089] Please see Figure 1 This is a schematic diagram of the hardware structure of a mobile terminal implementing various embodiments of this application. The mobile terminal 100 may include: an RF (Radio Frequency) unit 101, a WiFi module 102, an audio output unit 103, an A / V (Audio / Video) input unit 104, a sensor 105, a display unit 106, a user input unit 107, an interface unit 108, a memory 109, a processor 110, and a power supply 111, etc. Those skilled in the art will understand that... Figure 1 The mobile terminal structure shown does not constitute a limitation on the mobile terminal. The mobile terminal may include more or fewer components than shown, or combine certain components, or have different component arrangements.
[0090] The following is combined with Figure 1 A detailed introduction to each component of the mobile terminal:
[0091] The radio frequency unit 101 can be used for receiving and transmitting signals during information transmission or calls. Specifically, it receives downlink information from the base station and processes it with the processor 110; additionally, it transmits uplink data to the base station. Typically, the radio frequency unit 101 includes, but is not limited to, an antenna, at least one amplifier, a transceiver, a coupler, a low-noise amplifier, and a duplexer. Furthermore, the radio frequency unit 101 can also communicate wirelessly with networks and other devices. The aforementioned wireless communications may use any communication standard or protocol, including but not limited to GSM (Global System of Mobile communication), GPRS (General Packet Radio Service), CDMA2000 (Code Division Multiple Access 2000), WCDMA (Wideband Code Division Multiple Access), TD-SCDMA (Time Division-Synchronous Code Division Multiple Access), FDD-LTE (Frequency Division Duplexing-Long Term Evolution), TDD-LTE (Time Division Duplexing-Long Term Evolution), 5G, and 6G.
[0092] WiFi is a short-range wireless transmission technology. Mobile terminals, through the WiFi module 102, can help users send and receive emails, browse web pages, and access streaming media, providing users with wireless broadband internet access. Although Figure 1 WiFi module 102 is shown, but it is understood that it is not a necessary component of a mobile terminal and can be omitted as needed without changing the nature of the invention.
[0093] The audio output unit 103 can convert audio data received by the radio frequency unit 101 or the WiFi module 102 or stored in the memory 109 into audio signals and output them as sound when the mobile terminal 100 is in call signal receiving mode, call mode, recording mode, voice recognition mode, broadcast receiving mode, etc. Furthermore, the audio output unit 103 can also provide audio output related to specific functions performed by the mobile terminal 100 (e.g., call signal receiving sound, message receiving sound, etc.). The audio output unit 103 may include a speaker, a buzzer, etc.
[0094] The A / V input unit 104 is used to receive audio or video signals. The A / V input unit 104 may include a graphics processing unit (GPU) 1041 and a microphone 1042. The GPU 1041 processes image data of still images or videos acquired by an image capture device (such as a camera) in video capture mode or image capture mode. The processed image frames can be displayed on the display unit 106. The image frames processed by the GPU 1041 can be stored in the memory 109 (or other storage medium) or transmitted via the radio frequency unit 101 or the WiFi module 102. The microphone 1042 can receive sound (audio data) in operating modes such as telephone call mode, recording mode, and voice recognition mode, and can process such sound into audio data. The processed audio (voice) data can be converted into a format that can be transmitted to a mobile communication base station via the radio frequency unit 101 in telephone call mode. The microphone 1042 can implement various types of noise cancellation (or suppression) algorithms to eliminate (or suppress) noise or interference generated during the reception and transmission of audio signals.
[0095] The mobile terminal 100 also includes at least one sensor 105, such as a light sensor, a motion sensor, and other sensors. Optionally, the light sensor includes an ambient light sensor and a proximity sensor. Optionally, the ambient light sensor can adjust the brightness of the display panel 1061 according to the ambient light level, and the proximity sensor can turn off the display panel 1061 and / or backlight when the mobile terminal 100 is moved to the ear. As a type of motion sensor, an accelerometer sensor can detect the magnitude of acceleration in various directions (generally three axes), and can detect the magnitude and direction of gravity when stationary. It can be used for applications that recognize the phone's posture (such as landscape / portrait switching, related games, magnetometer posture calibration), vibration recognition related functions (such as pedometer, tapping), etc. Other sensors that may be configured in the phone, such as fingerprint sensors, pressure sensors, iris sensors, molecular sensors, gyroscopes, barometers, hygrometers, thermometers, and infrared sensors, will not be described in detail here.
[0096] The display unit 106 is used to display information input by the user or information provided to the user. The display unit 106 may include a display panel 1061, which may be configured in the form of a liquid crystal display (LCD), an organic light-emitting diode (OLED), or the like.
[0097] User input unit 107 can be used to receive input numerical or character information, and generate key signal inputs related to user settings and function control of the mobile terminal. Optionally, user input unit 107 may include touch panel 1071 and other input devices 1072. Touch panel 1071, also known as a touch screen, can collect touch operations performed by the user on or near it (such as operations performed by the user using a finger, stylus, or any suitable object or accessory on or near touch panel 1071), and drive corresponding connection devices according to a pre-set program. Touch panel 1071 may include a touch detection device and a touch controller. Optionally, the touch detection device detects the user's touch position and the signal generated by the touch operation, and transmits the signal to the touch controller; the touch controller receives touch information from the touch detection device, converts it into touch point coordinates, sends it to processor 110, and can receive and execute commands sent by processor 110. In addition, touch panel 1071 can be implemented using various types such as resistive, capacitive, infrared, and surface acoustic wave. In addition to the touch panel 1071, the user input unit 107 may also include other input devices 1072. Optionally, other input devices 1072 may include, but are not limited to, one or more of the following: physical keyboard, function keys (such as volume control buttons, power buttons, etc.), trackball, mouse, joystick, etc., without being specifically limited here.
[0098] Optionally, the touch panel 1071 may cover the display panel 1061. When the touch panel 1071 detects a touch operation on or near it, it transmits the information to the processor 110 to determine the type of touch event. Subsequently, the processor 110 provides corresponding visual output on the display panel 1061 based on the type of touch event. Although in Figure 1 In this embodiment, the touch panel 1071 and the display panel 1061 are two independent components to realize the input and output functions of the mobile terminal. However, in some embodiments, the touch panel 1071 and the display panel 1061 can be integrated to realize the input and output functions of the mobile terminal. The specific implementation is not limited here.
[0099] Interface unit 108 serves as an interface through which at least one external device can connect to mobile terminal 100. For example, the external device may include a wired or wireless headset port, an external power supply (or battery charger) port, a wired or wireless data port, a memory card port, a port for connecting a device with an identification module, an audio input / output (I / O) port, a video I / O port, a headphone port, and so on. Interface unit 108 may be used to receive input (e.g., data, power, etc.) from the external device and transmit the received input to one or more elements within mobile terminal 100, or it may be used to transmit data between mobile terminal 100 and the external device.
[0100] The memory 109 can be used to store software programs and various data. The memory 109 may primarily include a program storage area and a data storage area. Optionally, the program storage area may store the operating system, applications required for at least one function (such as sound playback, image playback, etc.), etc.; the data storage area may store data created based on the use of the mobile phone (such as audio data, phonebook, etc.). Furthermore, the memory 109 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device.
[0101] The processor 110 is the control center of the mobile terminal. It connects various parts of the mobile terminal via various interfaces and lines. By running or executing software programs and / or modules stored in the memory 109, and by calling data stored in the memory 109, it performs various functions and processes data of the mobile terminal, thereby providing overall monitoring of the mobile terminal. The processor 110 may include one or more processing units; preferably, the processor 110 may integrate an application processor and a modem processor. Optionally, the application processor mainly handles the operating system, user interface, and applications, while the modem processor mainly handles wireless communication. It is understood that the modem processor may not be integrated into the processor 110.
[0102] The mobile terminal 100 may also include a power supply 111 (such as a battery) that supplies power to various components. Preferably, the power supply 111 can be logically connected to the processor 110 through a power management system, thereby enabling functions such as charging, discharging, and power consumption management through the power management system.
[0103] although Figure 1 As not shown, the mobile terminal 100 may also include a Bluetooth module, etc., which will not be described in detail here.
[0104] To facilitate understanding of the embodiments of this application, the communication network system on which the mobile terminal of this application is based is described below.
[0105] Please see Figure 2 , Figure 2 This application provides a communication network system architecture diagram. The communication network system is an LTE system based on the universal mobile communication technology. The LTE system includes a UE (User Equipment) 201, an E-UTRAN (Evolved UMTS Terrestrial Radio Access Network) 202, an EPC (Evolved Packet Core) 203, and the operator's IP services 204, which are connected in sequence.
[0106] Optionally, UE201 can be the aforementioned terminal 100, which will not be described in detail here.
[0107] E-UTRAN202 includes eNodeB2021 and other eNodeB2022, etc. Optionally, eNodeB2021 can connect to other eNodeB2022 via backhaul (e.g., X2 interface), and eNodeB2021 connects to EPC203, providing access from UE201 to EPC203.
[0108] EPC203 may include MME (Mobility Management Entity) 2031, HSS (Home Subscriber Server) 2032, other MMEs 2033, SGW (Serving Gateway) 2034, PGW (Packet Data Network Gateway) 2035, and PCRF (Policy and Charging Rules Function) 2036, etc. Optionally, MME2031 is the control node that handles signaling between UE201 and EPC203, providing bearer and connection management. HSS2032 is used to provide registers to manage functions such as the Home Location Register (not shown in the figure) and stores user-specific information such as service characteristics and data rates. All user data can be sent through SGW2034. PGW2035 can provide UE 201 IP address allocation and other functions. PCRF2036 is the policy and charging control decision point for service data flow and IP bearer resources. It selects and provides available policy and charging control decisions for the policy and charging enforcement function unit (not shown in the figure).
[0109] IP services 204 may include the Internet, intranet, IMS (IP Multimedia Subsystem), or other IP services.
[0110] Although the above description uses the LTE system as an example, those skilled in the art should know that this application is not only applicable to the LTE system, but also to other wireless communication systems, such as GSM, CDMA2000, WCDMA, TD-SCDMA, 5G and future new network systems (such as 6G), etc., without limitation.
[0111] Figure 3 This is a schematic diagram of the hardware structure of a controller 140 provided in this application. The controller 140 includes a memory 1401 and a processor 1402. The memory 1401 is used to store program instructions, and the processor 1402 is used to call the program instructions in the memory 1401 to execute the steps performed by the controller in the first embodiment of the above method. The implementation principle and beneficial effects are similar, and will not be described again here.
[0112] Optionally, the controller further includes a communication interface 1403, which can be connected to the processor 1402 via a bus 1404. The processor 1402 can control the communication interface 1403 to implement the receiving and sending functions of the controller 140.
[0113] Figure 4 This application provides a schematic diagram of the hardware structure of a network node 150. The network node 150 includes a memory 1501 and a processor 1502. The memory 1501 is used to store program instructions, and the processor 1502 is used to call the program instructions in the memory 1501 to execute the steps performed by the first node in the first embodiment of the above method. The implementation principle and beneficial effects are similar, and will not be described again here.
[0114] Optionally, the controller further includes a communication interface 1503, which can be connected to the processor 1502 via a bus 1504. The processor 1502 can control the communication interface 1503 to implement the receiving and sending functions of the network node 150.
[0115] The integrated modules described above, implemented as software functional modules, can be stored in a computer-readable storage medium. These software functional modules, stored in a storage medium, include several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute some steps of the methods of the various embodiments of this application.
[0116] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product. A computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the flow or function according to the embodiments of this application is generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a storage medium or transmitted from one storage medium to another. For example, computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state disk, SSD), etc.
[0117] Based on the above-described mobile terminal hardware structure and communication network system, various embodiments of this application are proposed.
[0118] First Embodiment
[0119] Reference Figure 5 , Figure 5 This is a flowchart illustrating the image processing method according to the first embodiment. The image processing method of this application embodiment can be applied to a processing device, including step S1:
[0120] Step S1: Perform prediction processing on at least one image block based on the spatial features of at least one reference block.
[0121] In this embodiment, the processing device can be a smart terminal, such as a mobile phone or computer, or a server, such as a local server or a cloud server. This embodiment and this application primarily use a smart terminal as an example for illustration.
[0122] Optionally, the technical solution of this embodiment can be applied to fields such as image encoding and decoding, video encoding and decoding, hardware video encoding and decoding, dedicated circuit video encoding and decoding, and real-time video encoding and decoding.
[0123] Optionally, the processing device can pre-store various images and videos, and can select one image to be predicted from among the images as an image block, or cut the selected image and use the cut image block as the image block to be predicted. Alternatively, it can extract a frame from the video sequence as an image block, or cut the extracted frame to obtain an image block. Or, the processing device receives the input image or video, extracts a frame from the image or video as an image block, or cuts the extracted frame to obtain an image block. Alternatively, the processing device receives images or videos sent by other network devices, extracts a frame from the image or video as an image block, or cuts the extracted frame to obtain an image block. In this case, the processing device establishes a communication connection with the network device on the mobile communication system network side in advance, so that the network device can send images or videos to the terminal device through the communication connection, and the terminal device receives the images or videos.
[0124] Optionally, the technical solution of this embodiment can be performed for intra-frame prediction or inter-frame prediction. No limitation is imposed here.
[0125] Optionally, the spatial characteristics of the reference block may include at least one of the following: style characteristics of the reference block, gradient distribution, edge orientation histogram, color component correlation, and motion activity.
[0126] Optionally, gradient distributions such as those for VVC (Versatile Video Coding, a next-generation video coding standard jointly developed by the ITU Telecommunication Standardization Sector and the International Organization for Standardization / International Electrotechnical Commission Moving Picture Experts Group) and ALF (Adaptive Loop Filter) gradients can be calculated.
[0127] Optionally, the edge orientation histogram corresponds to AV1 (AOMedia Video 1, an open-source, royalty-free next-generation video coding standard developed by the Open Media Consortium) directional filtering.
[0128] Optionally, color component correlation covers VVC CCLM (Cross-Component Linear Model) chromaticity prediction. Optionally, motion activity is used for dynamic QP (quantization parameter) adjustment.
[0129] Optionally, one reference block may correspond to one spatial feature, or two or more reference blocks may correspond to one spatial feature, or two or more reference blocks may correspond to two or more spatial features. Optionally, each reference pixel in a reference block may correspond to at least one spatial feature.
[0130] Optionally, a spatial feature can be determined or obtained based on at least one reference pixel in at least one reference block.
[0131] Optionally, the spatial features of the reference block may include at least one spatial feature determined or obtained based on at least one reference pixel or a first reference intermediate element in the reference block.
[0132] Optionally, a spatial feature may be determined or obtained based on a reference pixel or a first reference intermediate element, or two or more spatial features may be determined or obtained based on a reference pixel or a first reference intermediate element, or a spatial feature may be determined or obtained based on two or more reference pixels or a first reference intermediate element, or two or more spatial features may be determined or obtained based on two or more reference pixels or a first reference intermediate element.
[0133] Optionally, the spatial features of the reference block may include the spatial features corresponding to all or part of the reference pixels or the first reference intermediate elements in the reference block. The spatial feature corresponding to a reference pixel or the first reference intermediate element may be determined or obtained based on at least one reference pixel or the first reference intermediate element in the reference block, such as based on the reference pixel or the first reference intermediate element and other reference pixels or the first reference intermediate elements in the reference block.
[0134] Optionally, the spatial characteristics of at least one reference block can be determined or obtained based on the statistical characteristic value of at least one reference pixel or the first reference intermediate element in the at least one reference block.
[0135] Optionally, the spatial features of at least one reference block can be determined or obtained by preprocessing the reference block and then determining or obtaining them based on the statistical feature values of at least one reference pixel or the first reference intermediate element / element in the preprocessed reference block.
[0136] Optionally, statistical feature values may include at least one of the following: mean, maximum, minimum, median, mode (the pixel value that appears most frequently), range (the difference between the maximum and minimum values), variance, standard deviation, etc., or may include values determined or obtained according to other rules or calculation methods, without limitation.
[0137] Optionally, the spatial characteristics of the reference block can be determined or obtained by sequentially calculating the statistical feature values of all reference pixels or the first reference intermediate element in the reference block, and then using these statistical feature values.
[0138] Optionally, the reference block is an image block that has already been predicted and / or reconstructed.
[0139] Optionally, the reference block for at least one image block may be other image blocks used to assist in the prediction processing of at least one image block.
[0140] Optionally, the reference block can be a rectangular block, a non-rectangular block, an L-shaped block, or a T-shaped block.
[0141] Optionally, the reference block may include multiple rows and / or columns of pixels, or it may include a row and / or a column of pixels.
[0142] Optionally, the reference block of at least one image block includes a reference prediction block and / or a reference reconstruction block of the reference block.
[0143] Optionally, the reference block is determined or obtained according to at least one of the following methods 1 to 6:
[0144] Method 1, at least one of the following: the top adjacent pixel, the top non-adjacent pixel, the left adjacent pixel, the left non-adjacent pixel, the top left adjacent pixel, and the top left non-adjacent pixel of the image block;
[0145] Alternatively, the pixel to be predicted in the image patch can be used as the first pixel.
[0146] Optionally, the adjacent pixel above can be a pixel in the same image block as the first pixel in the final segmentation, and / or the pixel is located above and adjacent to the first pixel.
[0147] Optionally, the non-adjacent pixel above can be a pixel that shares the same image block as the first pixel in the same image, which is ultimately divided into the same image blocks. Although the pixel is located above the first pixel, the two pixels are not adjacent.
[0148] Optionally, the left adjacent pixel can be a pixel that shares the same image block as the first pixel in the final image segmentation, and / or the pixel is located to the left of the first pixel and adjacent to the first pixel.
[0149] Optionally, the non-adjacent pixel on the left can be a pixel that shares the same image block as the first pixel in the final image segmentation, even though the pixel is located to the left of the first pixel, the two pixels are not adjacent.
[0150] Optionally, the upper left adjacent pixel can be a pixel in the same image block as the first pixel in the final division, and / or the pixel is located to the upper left of the first pixel and adjacent to the first pixel.
[0151] Optionally, the non-adjacent pixel in the upper left corner can be a pixel that shares the same image block as the first pixel in the final image segmentation. Although the pixel is located to the upper left of the first pixel, the pixel and the first pixel are not adjacent.
[0152] Optionally, the lower left adjacent pixel can be a pixel in the same image block as the first pixel in the final division, and / or the pixel is located to the lower left of the first pixel and adjacent to the first pixel.
[0153] Optionally, the non-adjacent pixel in the lower left corner can be a pixel that shares the same image block as the first pixel in the final image segmentation, and / or the pixel is located in the lower left corner of the first pixel and is not adjacent to the first pixel.
[0154] Optionally, the upper right adjacent pixel can be a pixel in the same image block as the first pixel in the final division, and / or the pixel is located to the upper right of the first pixel and adjacent to the first pixel.
[0155] Optionally, the non-adjacent pixel in the upper right corner can be a pixel that shares the same image block as the first pixel in the final image segmentation, and / or the pixel is located in the upper right corner of the first pixel and is not adjacent to the first pixel.
[0156] Optionally, at least one of the following can be a reconstructed pixel or a predicted pixel: the upper adjacent pixel, the upper non-adjacent pixel, the left adjacent pixel, the left non-adjacent pixel, the upper left adjacent pixel, the upper left non-adjacent pixel, the lower left adjacent pixel, the lower left non-adjacent pixel, the upper right adjacent pixel, and the upper right non-adjacent pixel.
[0157] Optionally, at least one of the following can be used as a pixel in the reference block: the upper adjacent pixel, the upper non-adjacent pixel, the left adjacent pixel, the left non-adjacent pixel, the upper left adjacent pixel, the upper left non-adjacent pixel, the lower left adjacent pixel, the lower left non-adjacent pixel, the upper right adjacent pixel, and the upper right non-adjacent pixel. Alternatively, at least one obtained pixel can be deduced or calculated to obtain the pixel in the reference block, and / or the reference block can be determined or obtained based on the pixel in the reference block.
[0158] In this embodiment, a reference block is determined or generated based on at least one of the following: the top adjacent pixel, the top non-adjacent pixel, the left adjacent pixel, the left non-adjacent pixel, the top left adjacent pixel, the top left non-adjacent pixel, the bottom left adjacent pixel, the bottom left non-adjacent pixel, the top right adjacent pixel, and the top right non-adjacent pixel. A prediction block is then determined or generated based on the derivative block corresponding to at least one reference block. Specifically, during prediction processing, the corresponding derivative block can be determined based on an accurate and effective reference block. Using this derivative block for prediction processing can improve the prediction effect of the prediction processing.
[0159] Method 2: At least one of the following: neighboring block, non-neighboring block, cross-component block, co-location block, temporal block, and default block, corresponding to the image block;
[0160] Optionally, the default block can be a pre-set block, such as a block with typical pixel characteristics pre-set by the encoder and / or decoder.
[0161] Optionally, the neighboring block can be a block adjacent to the image block, and / or a block that has already been predicted or reconstructed.
[0162] Optionally, a non-neighbor block can be a block that is not adjacent to an image block, and / or a block that has already been predicted or reconstructed.
[0163] Optionally, the co-position block can be an image block in the co-position image that has the same position and size as the image block. Optionally, the co-position image can be the image in the reference image that is closest to the current image in time.
[0164] Optionally, the temporal block can be a block that is distinguished in the time domain, such as an image block in the previous frame. For example, if there is video data containing three frames of images, the first frame is played in the first second, the second frame is played in the second second, and the third frame is played in the third second, if the image block to be predicted at the current moment (such as the block to be predicted) is an image block after the second frame is divided, then the temporal block can be determined to be the image block corresponding to it in the first frame.
[0165] Optionally, the cross-component block can be an image block in a different component than the current at least one image block. For example, if the current at least one image block is an image block of the Y component, then the cross-component block can be an image block of the U component and / or V component. Optionally, if the current at least one image block is an image block of the U component, then the cross-component block can be an image block of the Y component and / or V component. Optionally, if the current at least one image block is an image block of the V component, then the cross-component block can be an image block of the U component and / or Y component.
[0166] Optionally, at least one pixel can be obtained from at least one of the following: neighboring blocks, non-neighboring blocks, cross-component blocks, co-located blocks, temporal blocks, and default blocks corresponding to the image block. The obtained at least one pixel can be used as a pixel in the reference block, or the obtained at least one pixel can be derived or calculated to obtain the pixel in the reference block. The reference block is determined or obtained based on the pixels in the reference block.
[0167] Optionally, at least one of the following can be determined from at least one of the neighboring blocks, non-neighboring blocks, cross-component blocks, co-position blocks, temporal blocks, and default blocks corresponding to the image block: an upper adjacent pixel, an upper non-neighboring pixel, a left adjacent pixel, a left non-neighboring pixel, an upper-left adjacent pixel, and an upper-left non-neighboring pixel, and used as the pixels of the reference block. Alternatively, the corresponding adjacent regions and / or non-adjacent regions can be determined from at least one of the neighboring blocks, non-neighboring blocks, cross-component blocks, co-position blocks, temporal blocks, and default blocks corresponding to the image block, and at least one pixel can be selected from them as the pixels of the reference block. The reference block is determined or obtained based on the pixels in the reference block.
[0168] In this embodiment, a reference block is determined or generated based on at least one of the following: neighboring blocks, non-neighboring blocks, cross-component blocks, co-position blocks, temporal blocks, and default blocks corresponding to the image block. A prediction block is determined or generated based on the derived block corresponding to at least one reference block. Specifically, during prediction processing, the corresponding derived block can be determined based on an accurate and effective reference block. Using the derived block for prediction processing can improve the prediction effect of the prediction processing.
[0169] Method 3, at least one of the following: width, height, block size, and block area of the image block;
[0170] Optionally, at least one reference block is determined or obtained based on at least one of the width, height, block size, and block area of at least one image block.
[0171] Optionally, for example, if at least one of the width, height, block size, and block area of at least one image block is greater than a preset threshold, then at least one image block is selected as a reference block from the neighboring blocks, non-neighboring blocks, cross-component blocks, co-location blocks, temporal blocks, and default blocks corresponding to at least one image block.
[0172] Alternatively, at least one of the width, height, block size, and block area of at least one image block can be input into a neural network for determining a reference block, and the output can be a reference block.
[0173] In this embodiment, a reference block is determined or generated based on at least one of the width, height, block size, and block area of the image block, and a prediction block is determined or generated based on the derivative block corresponding to at least one reference block. Specifically, during prediction processing, the corresponding derivative block can be determined based on an accurate and effective reference block, and the prediction effect can be improved by using the derivative block for prediction processing.
[0174] Method 4: Candidate motion vectors or candidate block vectors of image blocks are used to determine or generate candidate blocks;
[0175] Optionally, the candidate motion vector or candidate block vector of the image block may include the motion vector or block vector corresponding to at least one of the neighboring blocks, non-neighboring blocks, cross-component blocks, co-position blocks, temporal blocks, and default blocks corresponding to the image block. It may also include the motion vector or block vector corresponding to at least one of the above adjacent pixels, above non-adjacent pixels, left adjacent pixels, left non-adjacent pixels, upper left adjacent pixels, and upper left non-adjacent pixels of the image block. It may also include the motion vector or block vector of the image block, etc. The following only uses the motion vector or block vector of the image block as an example.
[0176] Optionally, block vector calculation is performed on the image block, and candidate blocks corresponding to the image block are determined based on the block vector calculation results. For example, the pixels corresponding to the block vector calculation results are used as pixels in the candidate blocks.
[0177] Optionally, motion vector calculation is performed on the image patch, and candidate blocks corresponding to the image patch are determined based on the motion vector calculation results. For example, the pixels corresponding to the motion vector calculation results are used as pixels in the candidate blocks.
[0178] In this embodiment, a candidate block is determined or generated by the candidate motion vector or candidate block vector of the image block to determine or generate a reference block. Based on the derivative block corresponding to at least one reference block, a prediction block is determined or generated. Specifically, during prediction processing, the corresponding derivative block can be determined based on an accurate and effective reference block. Using the derivative block for prediction processing can improve the prediction effect of the prediction processing.
[0179] Method 5: If the first information of the image block satisfies the first condition, then the reference block is the first reference block;
[0180] Optionally, the first information of the image block can be at least one of the methods one through seven described above.
[0181] Optionally, the first condition can be any condition set by the user in advance, such as the width of the image block being greater than a preset width threshold, or the height of the image block being greater than a preset height threshold.
[0182] Optionally, the first reference block can be a pre-defined reference block, such as at least one of the following: neighboring block, non-neighboring block, cross-component block, co-location block, temporal block, candidate block, and default block corresponding to the image block.
[0183] Optionally, when the first information of an image block satisfies the first condition, a first reference block of a reference block can be determined, and a prediction block can be determined or generated based on the derived block corresponding to the first reference block of at least one image block.
[0184] In this embodiment, if the first information of an image block satisfies the first condition, then the reference block is the first reference block. The prediction block is determined or generated based on the derivative block corresponding to the first reference block. Specifically, during prediction processing, the corresponding derivative block can be determined based on the accurate and effective reference block. Using the derivative block for prediction processing can improve the prediction effect of the prediction processing.
[0185] Method 6: If the first information of the image block satisfies the second condition, then the reference block is the second reference block.
[0186] Optionally, the second condition can be any condition set by the user in advance, and / or can be different from the second condition. For example, if the first condition is that the width of the image block is greater than the preset width threshold, then the second condition can be that the width of the image block is less than the preset width threshold.
[0187] Optionally, the second reference block can be a pre-defined reference block, such as a block other than the first reference block among at least one of the following: neighboring blocks, non-neighboring blocks, cross-component blocks, co-location blocks, temporal blocks, candidate blocks, and default blocks corresponding to the image block.
[0188] In this embodiment, if the first information of an image block satisfies the second condition, then the reference block is the second reference block. The prediction block is determined or generated based on the derivative block corresponding to the second reference block. Thus, during prediction processing, the corresponding derivative block can be determined based on the accurate and effective reference block, and the prediction processing can be performed using the derivative block, thereby improving the prediction effect of the prediction processing.
[0189] Optionally, in step S1, at least one image block can be predicted based on the spatial characteristics of at least one reference block to determine or obtain the predicted block.
[0190] Optionally, the predicted block can be an image block that has undergone prediction processing.
[0191] Optionally, the prediction block may include at least one predicted pixel, or it may include pixels associated with at least one predicted pixel.
[0192] Optionally, pixel prediction can be performed on the encoding side as predicted pixels, or simply predicted pixels. Pixel reconstruction can be performed on the decoding side as predicted pixels, or simply predicted pixels.
[0193] Optionally, the processing device may be a decoding end, and if at the decoding end, the prediction block may be a decoded image block, and / or the prediction block may include pixel reconstruction.
[0194] Optionally, the processing device may be an encoding end, whereby the prediction block may be an image block that has undergone prediction processing, and / or the prediction block may include the prediction of pixels.
[0195] In this embodiment, by taking into account the spatial characteristics of the reference block of at least one image block when performing prediction processing on at least one image block, the prediction effect of the prediction processing can be improved.
[0196] Second Embodiment
[0197] Based on the first embodiment, a second embodiment is proposed.
[0198] In this embodiment, the spatial characteristics are determined or obtained according to at least one of the following methods one to eight:
[0199] Method 1: Statistical feature values of at least one reference pixel and at least one neighboring pixel in at least one reference block;
[0200] Optionally, the statistical feature value can be an indicator that quantifies the statistical attributes of pixels. It can be the feature mean (such as the mean of pixel values), maximum value, minimum value, variance, standard deviation, mode, median, etc. There are no restrictions here.
[0201] Optionally, a neighboring pixel can be a reference pixel that is adjacent to or in contact with the reference pixel within a certain range.
[0202] Optionally, the reference pixel can be a selected pixel in the reference block.
[0203] Optionally, statistical feature values of reference pixels and at least one neighboring pixel in at least one reference block can be statistically determined, and spatial features can be determined or obtained based on at least one statistical feature value. For example, the correspondence between each feature value and spatial features can be set in advance, and spatial features can be determined or obtained based on the correspondence.
[0204] Optionally, at least one spatial feature can be determined or obtained based on the statistical feature values of a reference pixel and at least one neighboring pixel in at least one reference block, and prediction processing can be performed on at least one image block based on the at least one spatial feature.
[0205] In this embodiment, the spatial domain features are determined or obtained based on the statistical feature values of the reference pixel and at least one neighboring pixel in at least one reference block, thereby ensuring the effectiveness of the spatial domain features. When using the spatial domain features to perform prediction processing on at least one image block, the key image block features can be focused on, which can improve the prediction effect on at least one image block.
[0206] Method 2: Statistical feature values of at least two non-neighboring pixels in at least one reference block;
[0207] Optionally, a non-neighbor pixel can be a pixel in the reference block that is relatively far from the reference pixel, and the pixel distance between the two is greater than a pixel distance threshold, which can be set in advance.
[0208] Optionally, statistical features can be calculated and statistically analyzed for at least two non-neighboring pixels in at least one reference block to determine or obtain corresponding statistical features. Then, spatial features can be determined or obtained based on the statistical features of at least two non-neighboring pixels in at least one reference block. For example, the correspondence between each feature value and spatial features can be set in advance, and spatial features can be determined or obtained based on the correspondence.
[0209] Optionally, at least one spatial feature can be determined or obtained based on the statistical feature values of a reference pixel and at least one non-neighbor pixel in at least one reference block, and prediction processing can be performed on at least one image block based on the at least one spatial feature.
[0210] In this embodiment, the spatial domain features are determined or obtained based on the statistical feature values of reference pixels and at least one non-neighboring pixel in at least one reference block, thereby ensuring the effectiveness of the spatial domain features. When using the spatial domain features to perform prediction processing on at least one image block, key image block features can be focused, which can improve the prediction effect on at least one image block.
[0211] Method 3: A reference pixel in at least one reference block, and at least one of the following statistical feature values: at least one above the reference pixel, at least one to the left of the reference pixel, at least one to the right of the reference pixel, and at least one below the reference pixel.
[0212] Optionally, the adjacent pixel can be a pixel in the reference block that is adjacent to the reference pixel or the first reference middle element. The adjacent pixels of the reference pixel or the first reference middle element can include at least one of the following: an upper adjacent pixel, at least one lower adjacent pixel, at least one right adjacent pixel, and at least one lower adjacent pixel.
[0213] Optionally, the spatial features of at least one reference block can be determined or obtained based on the statistical feature value of at least one of the following: a reference pixel or a first reference intermediate element in the at least one reference block, a reference pixel or a first reference intermediate element, a reference pixel or a first reference intermediate element, a reference pixel or a first reference intermediate element, a reference pixel or a first reference intermediate element, a reference pixel or a first reference intermediate element, a reference pixel or a first reference intermediate element, and ... a reference pixel or a first reference intermediate element, and a reference pixel or a first reference intermediate element, a reference pixel or a first reference intermediate element, a reference pixel or a first reference intermediate element, and a reference pixel or a first reference intermediate element, a reference pixel or a first reference intermediate element, a reference pixel or a
[0214] Optionally, for all or part of the reference pixels or first reference intermediate elements in at least one reference block, the statistical feature value of the reference pixel or first reference intermediate element in the at least one reference block and at least one of the following: at least one upper adjacent pixel, at least one left adjacent pixel, at least one right adjacent pixel, and at least one lower adjacent pixel of the reference pixel or first reference intermediate element can be determined, and at least one spatial feature of the at least one reference block can be determined or obtained based on the statistical feature value.
[0215] Optionally, the spatial features of at least one reference block can be determined or obtained based on the statistical feature values of a reference pixel or a first reference intermediate element in at least one reference block, and an upper adjacent pixel, a left adjacent pixel, a right adjacent pixel, and a lower adjacent pixel of the reference pixel or the first reference intermediate element, and the at least one image block can be predicted based on the spatial features of at least one reference block.
[0216] In this embodiment, spatial features are determined or obtained based on the statistical feature values of at least one of the following: a reference pixel in at least one reference block, at least one upper adjacent pixel, at least one left adjacent pixel, at least one right adjacent pixel, and at least one lower adjacent pixel. This ensures the effectiveness of the spatial features and allows for focusing on key image block features when using spatial features to perform prediction processing on at least one image block, thereby improving the prediction effect on at least one image block.
[0217] Method 4: The weighted statistical feature values of the reference pixel and the pixels adjacent to the reference pixel in at least one reference block;
[0218] Optionally, the weighted statistical feature value may include at least one of the following: reference pixel or first reference intermediate element in at least one reference block, weighted mean, weighted maximum value, weighted minimum value, weighted median, weighted mode (the pixel value that appears most frequently), weighted range (the difference between the maximum and minimum values), weighted variance, weighted standard deviation, etc., of the reference pixel or first reference intermediate element, or may include values determined or obtained according to other rules or calculation methods, without limitation.
[0219] Optionally, the spatial features of at least one reference block can be determined or obtained based on the weighted statistical feature value of at least one of the following: a reference pixel or a first reference intermediate element in the at least one reference block, at least one upper adjacent pixel, at least one left adjacent pixel, at least one right adjacent pixel, and at least one lower adjacent pixel of the reference pixel or the first reference intermediate element. The at least one image block can then be predicted based on the spatial features of the at least one reference block.
[0220] Optionally, the weights of the weighting can be preset.
[0221] Optionally, the weights can be obtained through pre-training / offline training of the neural network.
[0222] Optionally, the weights of different pixels can be the same or different.
[0223] Optionally, for all or part of the reference pixels or first reference intermediate elements in at least one reference block, a weighted statistical feature value can be determined for the reference pixel or first reference intermediate element in the at least one reference block, and at least one of the following: at least one upper adjacent pixel, at least one left adjacent pixel, at least one right adjacent pixel, and at least one lower adjacent pixel of the reference pixel or first reference intermediate element. At least one spatial feature of the at least one reference block can be determined or obtained based on the weighted statistical feature value.
[0224] Optionally, the spatial features of at least one reference block can be determined or obtained based on the weighted statistical feature values of a reference pixel or a first reference intermediate element in at least one reference block and an adjacent pixel above, a adjacent pixel to the left, a adjacent pixel to the right, and a adjacent pixel below the reference pixel or the first reference intermediate element, and the at least one image block can be predicted based on the spatial features of at least one reference block.
[0225] In this embodiment, spatial features are determined or obtained by using the weighted statistical feature values of the pixels adjacent to the reference pixels in at least one reference block, thereby ensuring the effectiveness of the spatial features. When using the spatial features to perform prediction processing on at least one image block, key image block features can be focused, which can improve the prediction effect on at least one image block.
[0226] Method 5: The statistical feature value of at least one pixel of the outermost layer of the window centered on the reference pixel, where the window size parameter is less than or equal to the reference block size parameter;
[0227] Optionally, the window can be a window centered on a reference pixel or a first reference middle element, and the window's size parameters can be set to be less than or equal to the size parameters of the reference block.
[0228] Optionally, the spatial features of at least one reference block can be determined or obtained by using the statistical feature value of at least one pixel at the outermost layer of a window centered on a reference pixel or a first reference intermediate element in at least one reference block, and the at least one image block can be predicted based on the spatial features of the at least one reference block.
[0229] Optionally, for all or part of the reference pixels or the first reference intermediate element in the at least one reference block, the statistical feature value of at least one pixel at the outermost layer of the window centered on the reference pixel or the first reference intermediate element can be determined, and at least one spatial feature of the at least one reference block can be determined or obtained based on the statistical feature value.
[0230] For example, if the size parameter of the reference block is 7×7 and the size parameter of the window is set to 5×5, the spatial characteristics of at least one reference block can be determined or obtained based on the statistical feature value of at least one pixel among the 16 pixels of the outermost layer of the 5×5 window centered on the reference pixel or the first reference middle element.
[0231] Optionally, the spatial features of at least one reference block can be determined or obtained by using the weighted statistical feature value of at least one pixel at the outermost layer of the window centered on the reference pixel or the first reference intermediate element in the at least one reference block, and the at least one image block can be predicted based on the spatial features of the at least one reference block.
[0232] In this embodiment, by using the statistical feature value of at least one pixel at the outermost layer of the window centered on the reference pixel, the size parameter of the window is less than or equal to the size parameter of the reference block to determine or obtain the spatial domain feature, thereby ensuring the effectiveness of the spatial domain feature. Thus, when using the spatial domain feature to perform prediction processing on at least one image block, the key image block features can be focused on, which can improve the prediction effect on at least one image block.
[0233] Method 6: Statistical feature values of eight pixels that are not adjacent to the reference pixel, and the eight pixels are not adjacent to each other;
[0234] Optionally, eight pixels can be determined from the pixels that are not adjacent to the reference pixel or the first reference intermediate element in at least one reference block. Based on the statistical feature values of the determined eight pixels, the spatial features of at least one reference block are determined or obtained, and the at least one image block is predicted based on the spatial features of at least one reference block.
[0235] Optionally, for all or part of the reference pixels or first reference intermediate elements in at least one reference block, statistical feature values that are not adjacent to the reference pixels or first reference intermediate elements in at least one reference block can be determined, and at least one spatial feature of at least one reference block can be determined or obtained based on the statistical feature values.
[0236] Optionally, for example, if the size of the reference block is 5×5, for the middle reference pixel or the first reference middle element in the reference block, 8 pixels can be determined from 16 pixels that are not adjacent to the reference pixel or the first reference middle element. Based on the statistical feature values of the determined 8 pixels, at least one spatial feature of at least one reference block can be determined or obtained.
[0237] Optionally, eight pixels can be determined from the pixels that are not adjacent to the reference pixel or the first reference intermediate element in at least one reference block. Based on the weighted statistical feature values of the determined eight pixels, the spatial features of at least one reference block are determined or obtained. The at least one image block is then subjected to prediction processing based on the spatial features of at least one reference block. The weighted weights of each pixel can be the same or different.
[0238] Optionally, the eight pixels that are not adjacent to the reference pixel or the first reference middle element can be mutually non-adjacent. For example, if the size of the reference block is 5×5, for the middle reference pixel or the first reference middle element in the reference block, eight non-adjacent pixels can be determined from the 16 non-adjacent pixels. Based on the statistical feature values of the eight determined non-adjacent pixels, at least one spatial feature of at least one reference block is determined or obtained, and prediction processing is performed on at least one image block based on the spatial feature of at least one reference block.
[0239] In this embodiment, spatial features are determined or obtained based on the statistical feature values of eight pixels that are not adjacent to the reference pixel, and the feature that the eight pixels are not adjacent to each other, thereby ensuring the effectiveness of the spatial features. Thus, when using the spatial features to perform prediction processing on at least one image block, key image block features can be focused on, which can improve the prediction effect on at least one image block.
[0240] Method 7: The statistical feature value of a reference sub-block containing at least one of the following: at least one reference pixel, at least one neighboring pixel, and at least one non-neighboring pixel;
[0241] Optionally, a reference block may include at least one reference sub-block, which may include at least one reference pixel, at least one neighboring pixel, and at least one non-neighboring pixel.
[0242] Optionally, each reference sub-block in a reference block may correspond to at least one spatial feature.
[0243] Optionally, a spatial feature of a reference block can be determined or obtained based on at least one reference sub-block within the reference block. Optionally, the spatial feature of a reference block may include at least one spatial feature determined or obtained based on at least one reference sub-block within the reference block.
[0244] Optionally, a reference sub-block may include at least one reference pixel or a first reference intermediate element and at least one neighboring pixel of the reference pixel or the first reference intermediate element; or, a reference sub-block may include at least one reference pixel or a first reference intermediate element and at least one non-neighboring pixel of the reference pixel or the first reference intermediate element; or, a reference sub-block may include at least one reference pixel or a first reference intermediate element, at least one neighboring pixel of the reference pixel or the first reference intermediate element, and at least one non-neighboring pixel of the reference pixel or the first reference intermediate element.
[0245] Optionally, a reference sub-block can be determined from the reference block, which includes at least one of at least one reference pixel or a first reference intermediate element, at least one neighboring pixel, and at least one non-neighboring pixel. The spatial features of at least one reference block are determined or obtained based on the statistical feature values of the reference sub-block, and the at least one image block is subjected to prediction processing based on the spatial features of the at least one reference block.
[0246] Optionally, for all or part of the reference pixels or first reference intermediate elements in at least one reference block, a reference sub-block can be determined, which includes at least one of the reference pixel or first reference intermediate element, at least one neighboring pixel of the reference pixel or first reference intermediate element, and at least one non-neighboring pixel of the reference pixel or first reference intermediate element. The spatial features of the at least one reference block are determined or obtained based on the statistical feature value of the reference sub-block, and the at least one image is predicted based on the spatial features of the at least one reference block.
[0247] Optionally, a reference sub-block can be determined from the reference block, which includes at least one of at least one reference pixel or a first reference intermediate element, at least one neighboring pixel, and at least one non-neighboring pixel. The spatial features of at least one reference block are determined or obtained based on the weighted statistical feature value of the reference sub-block, and the at least one image block is predicted based on the spatial features of the at least one reference block.
[0248] In this embodiment, spatial features are determined or obtained based on the statistical feature values of a reference sub-block that includes at least one of a reference pixel, at least one neighbor pixel, and at least one non-neighbor pixel. This ensures the effectiveness of the spatial features and allows for focusing on key image block features when using spatial features to perform prediction processing on at least one image block, thereby improving the prediction effect on at least one image block.
[0249] Method 8: At least one reference block contains statistical feature values of a reference region, including at least one reference pixel, at least one neighboring pixel, and at least one non-neighboring pixel.
[0250] Optionally, a reference block may include at least one reference region, and a reference region may include at least one reference sub-block. The concept of a reference sub-block can be referred to in Method 7 above.
[0251] Optionally, each reference region in a reference block may correspond to at least one spatial feature.
[0252] Optionally, a spatial feature of a reference block can be determined or obtained based on at least one reference region within the reference block. Alternatively, the spatial feature of a reference block may include at least one spatial feature determined or obtained based on at least one reference region within the reference block.
[0253] Optionally, a reference region can be determined from the reference block, which includes at least one of at least one reference pixel or a first reference intermediate element, at least one neighboring pixel, and at least one non-neighboring pixel. The spatial features of at least one reference block can be determined or obtained based on the statistical feature values of the reference region, and the at least one image block can be predicted based on the spatial features of the at least one reference block.
[0254] Optionally, for all or part of the reference pixels or first reference intermediate elements in at least one reference block, a reference region can be determined, which includes at least one of the reference pixels or first reference intermediate elements, at least one neighboring pixel of the reference pixel or first reference intermediate element, and at least one non-neighboring pixel of the reference pixel or first reference intermediate element. The spatial features of the at least one reference block can be determined or obtained based on the statistical feature value of the reference region, and the at least one image block can be predicted based on the spatial features of the at least one reference block.
[0255] Optionally, a reference region can be determined from the reference block, which includes at least one of at least one reference pixel or a first reference intermediate element, at least one neighboring pixel, and at least one non-neighboring pixel. The spatial features of at least one reference block can be determined or obtained based on the weighted statistical feature value of the reference region, and the at least one image block can be predicted based on the spatial features of the at least one reference block.
[0256] In this embodiment, spatial features are determined or obtained based on the statistical feature values of a reference region that includes at least one of at least one reference pixel, at least one neighboring pixel, and at least one non-neighboring pixel in at least one reference block. This ensures the effectiveness of the spatial features and allows for focusing on key image block features when using spatial features to perform prediction processing on at least one image block, thereby improving the prediction effect on at least one image block.
[0257] Third Embodiment
[0258] Based on any of the above embodiments, a third embodiment is proposed.
[0259] In this embodiment, step S1 includes at least one of the following steps a1 to a3:
[0260] Step a1: Based on at least one spatial feature and at least one lookup table, determine or obtain at least one spatial attention weight, and perform prediction processing on at least one image patch based on the at least one spatial attention weight;
[0261] Optionally, at least one spatial attention weight can be determined or obtained based on at least one spatial feature of at least one reference block and at least one lookup table, and prediction processing can be performed on at least one image block based on the at least one spatial attention weight.
[0262] Optionally, spatial attention weights can be spatial-level self-attention weights.
[0263] Optionally, spatial attention weights are weight coefficients that quantify the importance of different spatial locations (regions) in the feature map. By learning or computing the weight map, the model will allocate more attention to high-weight regions (enhancing features) and reduce attention to low-weight regions (suppressing irrelevant information), thereby focusing on the most discriminative regions in the image (such as the target subject, key textures, etc.).
[0264] Alternatively, lookup tables are a common method for accelerating computation in embedded systems. For a complex function or a series of calculations, if the output value is placed in a lookup table, then all that is needed afterward is to retrieve the value, without having to perform the computation process. Therefore, lookup tables are effective when the computation time is longer than the memory access time.
[0265] Alternatively, the lookup table can be used to approximate the computation of a neural network or neural network module, such as a 3x3 convolutional layer. The neural network or neural network module is pre-trained, and the lookup table is obtained based on the input-output mapping of the trained model.
[0266] Optionally, the lookup table may include at least one spatial feature, at least one spatial attention weight, and the correspondence between the two. Optionally, the lookup table may also include at least one spatial feature, at least one intermediate element, and the correspondence between the two. Optionally, the lookup table may also include at least one intermediate element, at least one spatial attention weight, and the correspondence between the two.
[0267] Optionally, at least one spatial attention weight can be determined or obtained in at least one lookup table based on at least one spatial feature of at least one image patch.
[0268] Optionally, after determining at least one spatial feature of at least one image patch, the at least one spatial feature can be input into at least one lookup table for searching. Since the lookup table includes the correspondence between at least one spatial feature and at least one spatial attention weight, the spatial attention weight that corresponds to the input at least one spatial feature can be found in the at least one lookup table and output. Then, the at least one image patch can be predicted based on the output at least one spatial attention weight.
[0269] Optionally, at least one spatial feature can be input into at least one lookup table to find at least one intermediate element, and at least one spatial attention weight can be determined or obtained based on the at least one intermediate element. For example, at least one intermediate element can be input into at least one neural network and / or at least one activation function to output at least one spatial attention weight.
[0270] Optionally, at least one intermediate element can be determined or obtained based on at least one spatial feature. For example, at least one spatial feature can be input into at least one neural network and / or at least one activation function to output at least one intermediate element, and the intermediate element can be input into at least one lookup table to obtain at least one spatial attention weight.
[0271] Optionally, the reference block may contain reference block features of at least one channel. Optionally, the spatial attention weights corresponding to the reference block features of at least one channel can be determined or obtained based on the spatial domain features of a reference block and at least one lookup table, and the image block can be predicted based on the spatial attention weights corresponding to the reference block features of one channel.
[0272] In this embodiment, at least one spatial attention weight is determined or obtained based on at least one spatial domain feature of at least one reference block and at least one lookup table. Prediction processing is performed on at least one image block based on the at least one spatial attention weight. The key image block features that need to be processed in the image block can be determined by the spatial attention weight, and the key image block features can be focused on during prediction processing, thereby improving the prediction effect of prediction processing on at least one image block.
[0273] Step a2: Based on at least one spatial feature and at least one neural network, determine or obtain at least one spatial attention weight, and perform prediction processing on at least one image patch based on the at least one spatial attention weight;
[0274] Optionally, at least one spatial attention weight can be determined or obtained based on at least one spatial domain feature of at least one reference block and at least one neural network, and prediction processing can be performed on at least one image block based on the at least one spatial attention weight.
[0275] Optionally, spatial attention weights can be spatial-level self-attention weights.
[0276] Optionally, the neural network in this embodiment can be a neural network based on fully connected layers, a neural network based on convolutional layers, a neural network based on Transformer, or a neural network based on a mixture of convolutional and fully connected layers and Transformer, etc.
[0277] Optionally, at least one spatial feature can be input into at least one neural network, and at least one spatial attention weight can be output.
[0278] Optionally, at least one spatial feature can be input into at least one neural network, outputting at least one intermediate element, and at least one spatial attention weight can be determined or obtained based on the at least one intermediate element. For example, at least one intermediate element can be input into at least one lookup table or at least one activation function, outputting at least one spatial attention weight.
[0279] Optionally, at least one intermediate element can be determined or obtained based on at least one spatial feature. For example, at least one spatial feature can be input into at least one lookup table or at least one activation function to output at least one intermediate element, and the intermediate element can be input into at least one neural network to output at least one spatial attention weight.
[0280] Optionally, at least one spatial feature can be input into at least one lookup table for searching, the intermediate element can be output, and the intermediate element can be input into at least one neural network to obtain at least one spatial attention weight.
[0281] Optionally, the lookup table may include at least one spatial feature, at least one intermediate element, and the correspondence between the two. Since the lookup table includes the correspondence between at least one spatial feature and at least one intermediate element, an intermediate element that corresponds to the input at least one spatial feature can be found in the at least one lookup table and output. The output at least one intermediate element is then input into at least one neural network to obtain at least one spatial attention weight. The at least one image patch is then used for prediction processing based on the at least one spatial attention weight.
[0282] Optionally, at least one spatial feature can be input into at least one neural network, the intermediate element can be output, the intermediate element can be input into at least one lookup table for searching, and at least one spatial attention weight can be output.
[0283] Optionally, the lookup table may include at least one intermediate element, at least one spatial attention weight, and the correspondence between the two. Since the lookup table includes the correspondence between at least one intermediate element and at least one spatial attention weight, after inputting at least one spatial feature into at least one neural network and outputting an intermediate element, the at least one intermediate element can be input into at least one lookup table to find at least one spatial attention weight that corresponds to the input at least one intermediate element, and prediction processing can be performed on at least one image patch based on the at least one spatial attention weight.
[0284] In this embodiment, by determining or obtaining at least one spatial attention weight based on at least one spatial domain feature of at least one reference block and at least one neural network, and performing prediction processing on at least one image block based on at least one spatial attention weight, the key image block features that need to be processed in the reference block can be determined through spatial attention, thereby focusing on the key image block features during prediction processing, which can improve the prediction effect of predicting at least one image block.
[0285] Step a3: Based on at least one spatial feature and at least one activation function, determine or obtain at least one spatial attention weight, and perform prediction processing on at least one image patch based on the at least one spatial attention weight.
[0286] Optionally, at least one spatial attention weight can be determined or obtained based on at least one spatial feature and at least one activation function of at least one reference block, and prediction processing can be performed on at least one image block based on the at least one spatial attention weight.
[0287] Optionally, spatial attention weights can be spatial-level self-attention weights.
[0288] Optionally, the activation function, also known as the activation function, often exists between the input and output layers of a neural network, and its purpose is to add some non-linear factors to the neural network.
[0289] Optionally, at least one spatial feature can be input into at least one activation function to output at least one spatial attention weight.
[0290] Optionally, at least one spatial feature can be input into at least one activation function to output at least one intermediate element, and at least one spatial attention weight can be determined or obtained based on the at least one intermediate element. For example, at least one intermediate element can be input into at least one lookup table or at least one neural network to output at least one spatial attention weight.
[0291] Optionally, at least one intermediate element can be determined or obtained based on at least one spatial feature. For example, at least one spatial feature can be input into at least one lookup table or at least one neural network to output at least one intermediate element, and the intermediate element can be input into at least one activation function to output at least one spatial attention weight.
[0292] Optionally, at least one spatial attention weight can be determined or obtained based on at least one spatial feature, at least one activation function, and at least one lookup table.
[0293] Optionally, at least one spatial feature can be input into at least one lookup table for searching, the intermediate element can be output, and the intermediate element can be input into at least one activation function to output at least one spatial attention weight.
[0294] Optionally, the lookup table may include at least one spatial feature, at least one intermediate element, and the correspondence between the two. Since the lookup table includes the correspondence between at least one spatial feature and at least one intermediate element, an intermediate element that corresponds to the input at least one spatial feature can be found in at least one lookup table and output. Then, the output at least one intermediate element is input into at least one activation function to obtain at least one spatial attention weight.
[0295] Optionally, at least one spatial feature can be input into at least one activation function, the intermediate element can be output, the intermediate element can be input into at least one lookup table for lookup, and at least one spatial attention weight can be output.
[0296] Optionally, the lookup table may include at least one intermediate element, at least one spatial attention weight, and the correspondence between the two. Since the lookup table includes the correspondence between at least one intermediate element and at least one spatial attention weight, after inputting at least one spatial feature into at least one activation function and outputting an intermediate element, the at least one intermediate element can be input into at least one lookup table to find at least one spatial attention weight that corresponds to the input at least one intermediate element.
[0297] Optionally, at least one spatial attention weight can be determined or obtained based on at least one spatial feature, at least one activation function, and at least one neural network.
[0298] Optionally, at least one spatial feature can be input into at least one activation function, the intermediate element can be output, and then the intermediate element can be input into at least one neural network to obtain at least one spatial attention weight. Alternatively, at least one spatial feature can be input into at least one neural network, the intermediate element can be output, and then the intermediate element can be input into at least one activation function to obtain at least one spatial attention weight.
[0299] Optionally, at least one spatial attention weight can be determined or obtained based on at least one spatial feature, at least one activation function, at least one neural network, and at least one lookup table.
[0300] In this embodiment, by determining or obtaining at least one spatial attention weight based on at least one spatial domain feature and at least one activation function of at least one reference block, and performing prediction processing on at least one image block based on at least one spatial attention weight, the key image block features that need to be processed in the reference block can be determined through spatial attention, thereby focusing on the key image block features during prediction processing, which can improve the prediction effect of prediction processing on at least one image block.
[0301] Fourth embodiment
[0302] Based on any of the above embodiments, a fourth embodiment is proposed.
[0303] In this embodiment, the image processing method further includes at least one of the following methods nine to twelve:
[0304] Method 9: At least two first intermediate elements in the intermediate element block corresponding to at least one reference block have the same spatial attention weight and / or the spatial attention weights corresponding to the reference pixels are different.
[0305] Optionally, the intermediate element block may contain at least one intermediate element, such as a first intermediate element. The first intermediate element can be an intermediate element. The intermediate element can be an element determined based on a reference pixel in the reference block. The intermediate element can be a reference pixel, and the intermediate pixel can be a pixel obtained by filtering, transforming, or downsampling the reference pixel.
[0306] Optionally, the intermediate element block can be an intermediate element block containing at least one intermediate element obtained by pre-filtering, transforming, or downsampling at least one reference pixel in at least one reference block.
[0307] Optionally, for at least two first intermediate elements in at least one intermediate element block, spatial attention weights corresponding to the at least two first intermediate elements can be determined or obtained, and the spatial attention weights corresponding to the at least two first intermediate elements can be different.
[0308] For at least two first intermediate elements in at least one intermediate element block, spatial attention weights corresponding to at least two first intermediate elements can be determined or obtained based on at least one spatial feature and at least one lookup table. The spatial attention weights corresponding to at least two first intermediate elements can be different. The at least one intermediate element block is then processed for prediction based on the spatial attention weights corresponding to at least two first intermediate elements.
[0309] For example, the size parameter of the intermediate element block is 3×3, and the spatial attention weights corresponding to at least two of the nine first intermediate elements in the intermediate element block can be different. For example, the spatial attention weights corresponding to the nine first intermediate elements in the image block are all different.
[0310] Optionally, the spatial attention weights corresponding to the first intermediate element in at least two rows or at least two columns of the intermediate element block can be different. For example, the spatial attention weight corresponding to the first intermediate element in the first row can be different from the spatial attention weight corresponding to the first intermediate element in the second row, or the spatial attention weight corresponding to the first intermediate element in each row can be different.
[0311] Optionally, for at least two first intermediate elements in at least one intermediate element block, spatial attention weights corresponding to the at least two first intermediate elements can be determined or obtained, and the spatial attention weights corresponding to the at least two first intermediate elements can be the same.
[0312] For example, the size parameter of the intermediate element block is 3×3, and the spatial attention weights corresponding to at least two of the nine first intermediate elements in the intermediate element block can be the same. For example, the spatial attention weights corresponding to all nine first intermediate elements in the intermediate element block can be the same.
[0313] Optionally, the spatial attention weights corresponding to the first intermediate element in at least two rows or at least two columns of the intermediate element block can be the same. For example, the spatial attention weight corresponding to the first intermediate element in the first row can be the same as the spatial attention weight corresponding to the first intermediate element in the second row, or the spatial attention weight corresponding to the first intermediate element in each row can be the same.
[0314] Optionally, for at least three first intermediate elements in at least one intermediate element block, spatial attention weights corresponding to the at least three first intermediate elements can be determined or obtained, wherein the spatial attention weights corresponding to at least two first intermediate elements can be different, and the spatial attention weights corresponding to at least two first intermediate elements can be the same.
[0315] For example, the size parameter of the intermediate element block is 3×3. Among the 9 first intermediate elements in the intermediate element block, at least two of the first intermediate elements can have the same spatial attention weight, and at least two of the first intermediate elements can have different spatial attention weights. For example, the spatial attention weights of the 4 first intermediate elements in the intermediate element block are all the same, and the spatial attention weights of the other 5 first intermediate elements are all different.
[0316] Optionally, the spatial attention weights corresponding to the first intermediate elements in at least two rows or at least two columns of the intermediate element block can be the same, or the spatial attention weights corresponding to the first intermediate elements in at least two rows or at least two columns can be different. For example, the spatial attention weights corresponding to the first intermediate elements in the first row and the first intermediate elements in the second row can be the same, or the spatial attention weights corresponding to the first intermediate elements in the second row and the first intermediate elements in the third row can be different.
[0317] Optionally, based on at least one spatial feature of at least one intermediate element block and at least one of at least one lookup table, at least one activation function and at least one neural network, spatial attention weights corresponding to at least two first intermediate elements in at least one intermediate element block can be determined or obtained. The spatial attention weights corresponding to at least two first intermediate elements can be different. Based on the different spatial attention weights corresponding to at least two first intermediate elements, prediction processing is performed on at least one image block.
[0318] In this embodiment, for at least two first intermediate elements in at least one intermediate element block, spatial attention weights corresponding to the at least two first intermediate elements can be determined or obtained. The spatial attention weights corresponding to the at least two first intermediate elements can be different and / or the same. Based on the at least one spatial attention weight, the key image block features that need to be processed in the intermediate element block can be determined by the different and / or the same spatial attention weights. In this way, the key image block features can be focused during the prediction process, which can improve the prediction effect of the prediction process for at least one image block.
[0319] Method 10: The spatial attention weight corresponding to at least one first intermediate element in the intermediate element block corresponding to at least one reference block is greater than the spatial attention weight corresponding to another reference pixel.
[0320] Optionally, for at least two first intermediate elements in at least one intermediate element block, spatial attention weights corresponding to at least two first intermediate elements can be determined or obtained, wherein the spatial attention weight corresponding to at least one first intermediate element is greater than the spatial attention weight corresponding to the other first intermediate element.
[0321] Optionally, for at least two first intermediate elements in at least one intermediate element block, spatial attention weights corresponding to at least two first intermediate elements can be determined or obtained based on at least one spatial feature and at least one lookup table, wherein the spatial attention weight corresponding to at least one first intermediate element is greater than the spatial attention weight corresponding to the other first intermediate element, and the at least one image block is predicted based on the spatial attention weights corresponding to at least two first intermediate elements.
[0322] For example, the size parameter of the intermediate element block is 3×3. The spatial attention weight corresponding to at least one of the nine first intermediate elements in the intermediate element block can be different from the spatial attention weight corresponding to the other first intermediate element. For example, the spatial attention weight corresponding to one first intermediate element in the intermediate element block is greater than the spatial attention weight corresponding to the other eight first intermediate elements.
[0323] Optionally, the spatial attention weight corresponding to the first intermediate element in at least one row or column of the intermediate element block may be greater than the spatial attention weight corresponding to the first intermediate element in another row or column. For example, the spatial attention weight corresponding to the first intermediate element in the first row may be greater than the spatial attention weight corresponding to the first intermediate element in the second row.
[0324] Optionally, spatial attention weights corresponding to at least two first intermediate elements in at least one intermediate element block can be determined or obtained based on at least one spatial feature and at least one of at least one lookup table, at least one activation function, and at least one neural network. Optionally, the spatial attention weight corresponding to at least one first intermediate element is greater than the spatial attention weight corresponding to another first intermediate element. Based on the spatial attention weights corresponding to at least two first intermediate elements, prediction processing is performed on at least one image block.
[0325] In this embodiment, for at least two first intermediate elements in at least one intermediate element block, spatial attention weights corresponding to the at least two first intermediate elements can be determined or obtained. The spatial attention weight corresponding to at least one first intermediate element is greater than the spatial attention weight corresponding to the other first intermediate element. Based on the at least one spatial attention weight, prediction processing is performed on at least one image block. Key image block features that need to be processed in the intermediate element block can be determined by spatial attention weights of different sizes. In this way, the key image block features can be focused on during prediction processing, which can improve the prediction effect of prediction processing on at least one image block.
[0326] Method 11: The spatial attention weight corresponding to the intermediate element region of at least one first intermediate element in the intermediate element block corresponding to at least one reference block is greater than the spatial attention weight corresponding to the intermediate element region of at least one other first intermediate element.
[0327] Optionally, the intermediate element block includes at least one intermediate element region, and the intermediate element region includes at least one first intermediate element. Optionally, an intermediate element region may correspond to at least one spatial attention weight, or at least one intermediate element region may correspond to one spatial attention weight.
[0328] Optionally, for at least two intermediate element regions in at least one intermediate element block, spatial attention weights corresponding to the at least two intermediate element regions can be determined or obtained, wherein the spatial attention weight corresponding to the intermediate element region containing at least one first intermediate element is greater than the spatial attention weight corresponding to the intermediate element region containing at least one other first intermediate element.
[0329] Optionally, for at least two intermediate element regions in at least one intermediate element block, spatial attention weights corresponding to the at least two intermediate element regions can be determined or obtained based on at least one spatial feature and at least one lookup table. The spatial attention weight corresponding to the intermediate element region containing at least one first intermediate element is greater than the spatial attention weight corresponding to the intermediate element region containing at least one other first intermediate element. The at least one image block is then subjected to prediction processing based on the spatial attention weights corresponding to the at least two intermediate element regions.
[0330] For example, the size parameter of the middle element block is 3×3. The middle element block contains two middle element regions. One middle element region contains the first middle element of the first row and the second row, and the other middle element region contains the first middle element of the third row. The spatial attention weight corresponding to the middle element region containing the first middle element of the first row and the second row can be greater than the spatial attention weight corresponding to the middle element region containing the first middle element of the third row.
[0331] Optionally, spatial attention weights corresponding to at least two intermediate element regions in at least one intermediate element block can be determined or obtained based on at least one spatial feature and at least one of at least one lookup table, at least one activation function, and at least one neural network. The spatial attention weights corresponding to the intermediate element regions containing at least one first intermediate element are greater than the spatial attention weights corresponding to the intermediate element regions containing at least one other first intermediate element. Based on the spatial attention weights corresponding to the at least two intermediate element regions, prediction processing is performed on at least one image block.
[0332] In this embodiment, for at least two intermediate element regions in at least one intermediate element block, spatial attention weights corresponding to at least two intermediate element regions can be determined or obtained. The spatial attention weight corresponding to the intermediate element region of at least one first intermediate element is greater than the spatial attention weight corresponding to the intermediate element region of at least one other first intermediate element. Based on the at least one spatial attention weight, prediction processing is performed on at least one image block. By using spatial attention weights of different sizes, key image block features that need to be processed in the intermediate element block can be determined. In this way, the key image block features can be focused during prediction processing, which can improve the prediction effect of prediction processing on at least one intermediate element block.
[0333] Method 12: The spatial attention weight corresponding to the intermediate element sub-block containing at least one first intermediate element in the intermediate element block corresponding to at least one reference block is greater than the spatial attention weight corresponding to the intermediate element sub-block containing at least one other first intermediate element.
[0334] Optionally, the intermediate element block includes at least one intermediate element region, the intermediate element region includes at least one intermediate element sub-block, and the intermediate element sub-block includes at least one first intermediate element.
[0335] Optionally, at least one intermediate element sub-block may correspond to at least one spatial attention weight, or at least one intermediate element sub-block may correspond to one spatial attention weight.
[0336] Optionally, for at least two intermediate element sub-blocks in at least one intermediate element block, spatial attention weights corresponding to the at least two intermediate element sub-blocks can be determined or obtained, wherein the spatial attention weight corresponding to the intermediate element sub-block containing at least one first intermediate element is greater than the spatial attention weight corresponding to the intermediate element sub-block containing at least one other first intermediate element.
[0337] Optionally, for at least two intermediate element sub-blocks in at least one intermediate element block, spatial attention weights corresponding to the at least two intermediate element sub-blocks can be determined or obtained based on at least one spatial feature and at least one lookup table. The spatial attention weight corresponding to the intermediate element sub-block containing at least one first intermediate element is greater than the spatial attention weight corresponding to the intermediate element sub-block containing at least one other first intermediate element. The at least one image block is then subjected to prediction processing based on the spatial attention weights corresponding to the at least two intermediate element sub-blocks.
[0338] For example, if the size parameter of the intermediate element block is 5×5, and the intermediate element block contains a 3×3 intermediate element region, the intermediate element region contains two intermediate element sub-blocks. One intermediate element sub-block contains the first intermediate element of the first row and the second row of the intermediate element region, and the other intermediate element sub-block contains the first intermediate element of the third row of the intermediate element region. The spatial attention weight corresponding to the intermediate element sub-block containing the first intermediate element of the first row and the second row can be greater than the spatial attention weight corresponding to the intermediate element sub-block containing the first intermediate element of the third row.
[0339] Optionally, spatial attention weights corresponding to at least two intermediate element sub-blocks in at least one intermediate element block can be determined or obtained based on at least one spatial feature of at least one intermediate element block and at least one item of at least one lookup table, at least one activation function, and at least one neural network. The spatial attention weights corresponding to the intermediate element sub-blocks containing at least one first intermediate element are greater than the spatial attention weights corresponding to the intermediate element sub-blocks containing at least one other first intermediate element. Based on the spatial attention weights corresponding to at least two intermediate element sub-blocks, prediction processing is performed on at least one image block.
[0340] In this embodiment, for at least two intermediate element sub-blocks in at least one intermediate element block, spatial attention weights corresponding to the at least two intermediate element sub-blocks can be determined or obtained. The spatial attention weights corresponding to the intermediate element sub-blocks containing at least one first intermediate element are greater than the spatial attention weights corresponding to the intermediate element sub-blocks containing at least one other first intermediate element. Based on the at least one spatial attention weight, prediction processing is performed on at least one image block. By using spatial attention weights of different sizes, the key image block features that need to be processed in the intermediate element block can be determined. In this way, the key image block features can be focused during prediction processing, which can improve the prediction effect of prediction processing on at least one image block.
[0341] Fifth Embodiment
[0342] Based on any of the above embodiments, a fifth embodiment is proposed.
[0343] In this embodiment, at least one intermediate element block is determined or obtained based on at least one neural network, a lookup table structure, at least one item in the lookup table, and at least one reference block.
[0344] Alternatively, the neural network can be a neural network module, such as a 3x3 convolutional layer.
[0345] Optionally, when performing prediction processing on at least one image block, at least one reference block can be input into a neural network for processing to determine or obtain at least one intermediate element block, and then subsequent prediction processing can be performed based on the at least one intermediate element block.
[0346] Optionally, a lookup index corresponding to at least one reference block can be determined, and a lookup table can be searched based on the index. At least one intermediate element block can be determined or obtained based on the lookup result, and subsequent prediction processing can be performed based on the at least one intermediate element block.
[0347] Optionally, at least one reference block can be processed based on at least one lookup table structure to determine or obtain at least one intermediate element block, and then subsequent prediction processing can be performed based on at least one intermediate element block.
[0348] Optionally, at least one reference block can be preprocessed, such as by a simple transformation, to determine or obtain at least one intermediate element block, and then subsequent prediction processing can be performed based on the at least one intermediate element block.
[0349] Optionally, the reference block contains at least one reference pixel.
[0350] Optionally, the first reference intermediate element can be an intermediate element, and the two can be at different stages. Optionally, the intermediate element can refer to the above description.
[0351] Optionally, the lookup table structure includes at least one of the following methods thirteen through twenty-nine:
[0352] Method 13 requires at least one lookup table;
[0353] Method fourteen, requiring at least one neural network module;
[0354] Optionally, the neural network module can be a type of neural network, such as a neural network containing only 3x3 convolutional layers, or a neural network module having at least one convolutional layer.
[0355] Optionally, the lookup table structure may include at least one lookup table, at least one neural network module, or both at least one lookup table and at least one neural network module.
[0356] Optionally, the index corresponding to at least one reference pixel or the first reference intermediate element can be determined, and the index can be input into at least one lookup table in the lookup table structure for searching. The predicted block pixel or the predicted intermediate element can be determined or obtained based on the search result. The at least one reference pixel or the first reference intermediate element can be input into the neural network module, and the predicted pixel or the predicted intermediate element can be output.
[0357] Optionally, the predicted pixel can be a pixel obtained by predicting the pixel to be predicted in the image patch.
[0358] Optionally, an index corresponding to at least one reference pixel or a first reference intermediate element can be determined, and the index can be input into at least one lookup table in the lookup table structure for searching. The search result is then input into the neural network module for processing. Based on the output of the neural network module, the predicted pixel after prediction processing can be determined or obtained. For example, the predicted pixel can be directly output, or a prediction intermediate element can be output and then input into at least one lookup table structure (such as a lookup table or a neural network module) for processing to finally obtain the predicted pixel. The prediction intermediate element can be an intermediate output generated during the prediction processing, such as a lookup table index, an intermediate result after downsampling, pixel statistical feature values, an autocorrelation matrix, or feature values output by the neural network module, attention weights, etc.
[0359] Optionally, at least one reference pixel or a first reference intermediate element can be input into the neural network module. At least one index can be determined based on the output of the neural network module, and the at least one index can be input into at least one lookup table for searching. The predicted pixel can be determined or obtained based on the search result. For example, the search result is a predicted pixel, a feature value, or a prediction pattern. The pixel to be predicted is processed based on the feature value or prediction pattern. Alternatively, it can be a prediction intermediate element. The prediction intermediate element is combined with the structure of at least one lookup table for further processing to finally obtain the predicted pixel.
[0360] In this embodiment, when the lookup table structure is at least one lookup table or at least one neural network module, at least one lookup table or at least one neural network module can be used to perform prediction processing on at least one pixel to be predicted. In this way, the lookup table or neural network module can be selected for prediction processing according to different scenarios to improve the prediction effect of prediction processing, thereby supporting the improvement of video encoding and / or decoding efficiency.
[0361] Method 15: There are at least two search branches, and at least one search branch has the same input as the other search branch;
[0362] Optionally, the lookup table structure may include at least two lookup branches, and each lookup branch may include at least one lookup table. The at least one lookup table included in each of the at least two lookup branches may be the same or different, and there is no restriction on this.
[0363] Optionally, at least two search branches may have the same input, such as at least one of the following: input pixel value, input number of pixels.
[0364] Optionally, the inputs to at least one lookup table in at least one lookup table in another lookup branch can be the same, for example, both can come from the same prediction intermediate element block (e.g., the block to be predicted, or a reference image block determined or obtained based on the block to be predicted).
[0365] For example, the input of at least one lookup table in at least one lookup branch is two pixels, and the input of at least one lookup table in another lookup branch can also be two pixels. The two input pixels can be the pixels to be predicted, or intermediate elements of prediction determined or obtained based on the pixels to be predicted, etc.
[0366] In this embodiment, when the inputs of at least two lookup branches in the lookup table structure are the same, multiple prediction processes can be performed based on at least one pixel to be predicted combined with at least two lookup branches in the lookup table structure that have the same input, thereby improving the prediction effect.
[0367] Method 16: At least two search branches, where the input of at least one search branch is determined or obtained based on the output of the other search branch;
[0368] Optionally, there are at least two search branches in the lookup table structure, and the inputs and outputs of the two search branches can have a certain correlation.
[0369] Optionally, the input of at least one search branch can be determined or obtained based on the output of another search branch. For example, the output of at least one search branch can be used as the input of another search branch, or the output range of at least one search branch can be used as the input range of another search branch, and so on.
[0370] Optionally, the input of at least one lookup table in at least one search branch can be determined or obtained based on the output of at least one lookup table in another search branch. For example, the output of at least one lookup table in at least one search branch can be used as the input of at least one lookup table in another search branch, or the output range of at least one lookup table in at least one search branch can be used as the input range of at least one lookup table in another search branch, and so on.
[0371] Optionally, the input of at least one lookup branch can be determined or obtained based on at least one reference pixel or a first reference intermediate element. For example, the at least one reference pixel or the first reference intermediate element can be used as the input of at least one lookup branch for prediction processing, or the index corresponding to at least one reference pixel or the first reference intermediate element can be used as the input of at least one lookup branch for prediction processing, and then the output of the at least one lookup branch can be used as the input of another lookup branch for prediction processing, until the predicted pixel is finally obtained.
[0372] In this embodiment, when the input of at least one lookup branch in the lookup table structure is determined or obtained based on the output of the other lookup branch, prediction processing can be performed on at least one pixel to be predicted in sequence according to these at least two lookup branches, which can improve the prediction effect.
[0373] Method 17: At least two search branches, with at least two search branches located in at least one channel corresponding to at least one reference pixel or the first reference intermediate element;
[0374] Alternatively, the channels in this embodiment can be used for lookup tables.
[0375] Optionally, a channel is a component of the feature map of an image patch in the depth dimension, used to describe the feature representation of the number of features in a specific dimension. Each channel represents a certain feature (such as texture, edge, derived distribution, etc.) extracted from the image patch (such as a prediction patch or reconstructed patch). For example, one channel might detect horizontal edges, while another channel might detect vertical edges, etc.
[0376] Optionally, at least one reference pixel may correspond to multiple channels.
[0377] Optionally, at least one first prediction intermediate element may correspond to multiple channels.
[0378] Optionally, in the lookup table structure, there are at least two lookup branches, and the at least two lookup branches are located in the same channel.
[0379] Alternatively, at least two lookup branches can be used as a predictor-like device to predict at least one reference pixel or a first reference intermediate element in the same channel.
[0380] Optionally, when processing the pixel features (such as pixel value, pixel position, etc.) of at least one reference pixel or the first reference intermediate element in the same channel, at least two search branches can be processed in parallel or in sequence, without any restriction.
[0381] Optionally, prediction processing can be performed on at least one reference pixel or a first reference intermediate element in the same channel based on at least one lookup table in each of at least two lookup branches.
[0382] In this embodiment, by using at least two lookup branches in the lookup table structure, and with at least two lookup branches located in at least one channel corresponding to at least one reference pixel or the first reference intermediate element, the pixel features (such as pixel value or pixel position) of at least one reference pixel or the first reference intermediate element in the same channel can be predicted using at least two lookup branches. This demonstrates the implementation of multiple prediction processing of pixel features in the same channel, thereby improving the prediction effect.
[0383] Method 18: At least two search branches, where the search table with at least two search branches is searched in parallel and / or serial manner.
[0384] Optionally, the lookup table structure includes at least two lookup branches, and the lookup methods of the at least two lookup branches are parallel lookup and / or serial lookup.
[0385] Optionally, the input elements corresponding to the lookup tables in at least two lookup branches can be determined simultaneously based on at least one reference pixel or a first reference intermediate element, and then input into their respective lookup tables for parallel lookup to obtain at least one prediction intermediate element. Then, subsequent prediction processing methods are used to perform prediction processing to obtain the prediction pixel.
[0386] Optionally, the input element corresponding to the lookup table in at least one lookup branch can be determined or obtained based on at least one reference pixel or a first reference intermediate element, and then input into the lookup table in that lookup branch for searching. After the search is completed, the search operation in the lookup table of another lookup branch is performed, that is, a serial search is performed until the predicted pixel or predicted intermediate element is finally determined or obtained.
[0387] In this embodiment, by using parallel search and / or serial search in at least two search branches of the lookup table structure, different search methods can be selected according to different scenarios when performing prediction processing on at least one pixel to be predicted using at least two search branches of the lookup table structure, thereby improving the prediction effect.
[0388] Method 19: At least two lookup tables, where the input of at least one lookup table is determined or obtained based on the type, size parameters, and / or input range of the other lookup table;
[0389] Optionally, there can be multiple types of lookup tables. Multiple lookup tables can be of the same type or different types, such as a predicted pixel lookup table, a homography matrix lookup table, an eigenvalue lookup table, etc.
[0390] Optionally, the size parameters of the lookup table may include the width, height, perimeter, and area of the lookup table.
[0391] Optionally, the input range of the lookup table can be set in advance, and the input ranges of multiple lookup tables can be the same or different.
[0392] Optionally, the lookup table structure includes at least two lookup tables. The input of one lookup table can be determined based on the type of the other lookup table. For example, if two consecutive lookup tables are of type predictive pixel lookup tables, the input range of one predictive pixel lookup table is greater than or equal to the input range of the other predictive pixel lookup table.
[0393] Optionally, the lookup table structure may contain at least two lookup tables, and the input of one lookup table may be determined based on the size parameters of the other lookup table. For example, the input range of the lookup table with the larger size parameter may be greater than or equal to the input range of the lookup table with the smaller size parameter.
[0394] Optionally, the lookup table structure may contain at least two lookup tables, and the input of one lookup table may be determined based on the input range of the other lookup table, such as the inputs of the two lookup tables being identical.
[0395] In this embodiment, by using at least two lookup tables in the lookup table structure, where the input of at least one lookup table is determined or obtained based on the type, size parameters, and / or input range of the other lookup table, multiple prediction processes can be performed on at least one pixel to be predicted using at least two lookup tables, thereby improving the prediction effect.
[0396] Method 20: At least two lookup tables, located in at least one channel corresponding to at least one reference pixel or the first reference intermediate element;
[0397] Optionally, at least one pixel to be predicted corresponds to multiple channels, and at least one first prediction intermediate element corresponds to multiple channels.
[0398] Optionally, the lookup table structure may include at least two lookup tables, which may be located in at least one channel of at least one pixel to be predicted, i.e., at least one lookup table exists in each channel of at least one pixel to be predicted, or at least two lookup tables exist in the same channel of at least one pixel to be predicted.
[0399] Optionally, the at least two lookup tables may be located in at least one channel of at least one first predicted intermediate element, that is, there is at least one lookup table in each channel of at least one first predicted intermediate element, or there may be at least two lookup tables in the same channel of at least one first predicted intermediate element.
[0400] In this embodiment, by using at least two lookup tables in the lookup table structure, with the at least two lookup tables located in at least one channel corresponding to at least one reference pixel or the first reference intermediate element, it is possible to perform multiple prediction processes on at least one reference pixel or the first reference intermediate element in the same channel using at least two lookup tables, thereby improving the prediction effect.
[0401] Method 21: At least two lookup tables, where the lookup method for at least two lookup tables within the same channel is parallel lookup and / or serial lookup;
[0402] Optionally, in the lookup table structure, the lookup method for at least two lookup tables is parallel lookup. When performing a parallel lookup on at least two lookup tables, the input of the two lookup tables comes from the same block (i.e., the reference block or the intermediate element block). By performing a one-time scan on the reference block or the intermediate element block, the indexes of at least two lookup tables can be obtained simultaneously, and then the at least two lookup tables can be performed in parallel based on the indexes of at least two lookup tables.
[0403] Optionally, the intermediate element block can be a reference block containing the predicted intermediate elements, which can be described as described above and will not be repeated here.
[0404] Optionally, in the lookup table structure, when at least two lookup tables are searched in parallel, the inputs and outputs of the two lookup tables are not dependent on each other. That is, the inputs of at least two lookup tables can come from different blocks or from the same block; at least two lookup tables can be searched simultaneously by multiple tasks / processes / hardware.
[0405] Optionally, in the lookup table structure, the lookup method for at least two lookup tables is serial lookup. When performing a serial lookup on at least two lookup tables, the input and output of the two lookup tables are dependent on each other. The operation of one lookup table must be performed first before the operation of the second lookup table is completed.
[0406] Optionally, the lookup table structure may include at least two lookup tables, and at least two lookup tables exist in the same channel corresponding to at least one pixel to be predicted, wherein the lookup method of the at least two lookup tables is parallel lookup and / or serial lookup.
[0407] Optionally, at least two lookup tables exist in the same channel corresponding to at least one first predicted intermediate element, and the lookup methods of the at least two lookup tables are parallel lookup and / or serial lookup.
[0408] Optionally, when processing the pixel to be predicted based on the lookup table structure, at least two lookup tables in the same channel can be used to perform parallel and / or serial lookups on the index corresponding to at least one reference pixel or the first reference intermediate element until the predicted pixel is finally obtained.
[0409] In this embodiment, by using parallel search and / or serial search in at least two search tables within the same channel in the lookup table structure, different search methods can be selected according to different scenarios when performing prediction processing on at least one pixel to be predicted in the same channel using at least two search branches of the lookup table structure, thereby improving the prediction effect.
[0410] Method 22: At least two lookup tables, and the lookup tables in at least two channels are searched in parallel and / or serial ways;
[0411] Optionally, the lookup table structure includes at least two lookup tables, with at least two channels corresponding to at least one pixel to be predicted containing lookup tables, and the lookup methods of the lookup tables in these two channels are parallel lookup and / or serial lookup.
[0412] Optionally, at least two channels corresponding to at least one first predicted intermediate element have lookup tables, and the lookup tables in these two channels are searched in parallel and / or serial.
[0413] Optionally, when processing the pixel to be predicted based on the lookup table structure, the lookup tables in at least two channels can be used to perform parallel searches of the index corresponding to the pixel to be predicted, or a serial search can be used to sequentially call the lookup tables in at least two channels until the predicted pixel is finally obtained.
[0414] Optionally, when processing the first predicted intermediate element according to the lookup table structure, the lookup tables in at least two channels can be used simultaneously to perform parallel lookups on the index corresponding to the first predicted intermediate element, or a serial lookup method can be used to sequentially call the lookup tables in at least two channels until the predicted pixel is finally obtained.
[0415] In this embodiment, by using parallel search and / or serial search in at least two search tables of the lookup table structure for at least two channels, different search methods can be selected according to different scenarios when performing prediction processing on at least one pixel to be predicted in different channels using at least two search branches of the lookup table structure, thereby improving the prediction effect.
[0416] Method 23: At least one search branch and at least one neural network module, wherein the input of the at least one neural network module is the same as the input of the at least one search branch;
[0417] Optionally, the neural network module may include a neural network, such as at least one convolutional layer, a 3x3 convolutional layer, etc.
[0418] Optionally, the input to at least one lookup branch in the lookup table structure is the same as the input to at least one neural network module.
[0419] Optionally, at least one reference pixel or a first reference intermediate element can be simultaneously input into at least one search branch and at least one neural network module for parallel processing. The output results of the parallel processing of at least one search branch and at least one neural network module can be fused or weighted before subsequent prediction processing is performed until the predicted pixel is finally obtained.
[0420] In this embodiment, by including at least one lookup branch and at least one neural network module with the same input in the lookup table structure, it is possible to use at least one lookup branch of the lookup table and at least one neural network module to jointly perform prediction processing on at least one reference pixel or the first reference intermediate element, and combine the advantages of both the lookup table in the lookup branch and the neural network in the neural network module to perform prediction processing, so as to improve the prediction effect.
[0421] Method 24: at least one search branch and at least one neural network module, wherein the input of the at least one neural network module is determined or obtained based on the output of the at least one search branch;
[0422] Optionally, the input of at least one neural network module in the lookup table structure can be determined or obtained based on the output of at least one lookup branch, and the input of at least one neural network module can be determined or obtained based on the output of at least one lookup table.
[0423] Optionally, the index corresponding to at least one reference pixel or the first reference intermediate element can be input into the lookup table in at least one lookup branch to perform a lookup, obtain the output of at least one lookup branch, and use the output of at least one lookup branch as the input of at least one neural network module (or the output of at least one lookup branch can be transformed, such as weighted processing, to obtain the input of at least one neural network module), and input into at least one neural network module for processing until all prediction processing steps are completed and the predicted pixel is obtained.
[0424] In this embodiment, by including at least one lookup branch and at least one neural network module in the lookup table structure, and determining or obtaining the input of the at least one neural network module based on the output of the at least one lookup branch, it is possible to use at least one lookup branch of the lookup table and at least one neural network module to jointly perform prediction processing on at least one reference pixel or first reference intermediate element. The prediction processing combines the advantages of both the lookup table in the lookup branch and the neural network in the neural network module to improve the prediction effect.
[0425] Method 25: at least one lookup table and at least one neural network module, wherein the input of the at least one neural network module is determined or obtained based on the type, size parameters and / or input range of the at least one lookup table;
[0426] Optionally, the type, size parameters, and / or input range of at least one lookup table can be determined in accordance with the above method ten.
[0427] Optionally, the lookup table structure may include at least one lookup table and at least one neural network module. At least one reference pixel or a first reference intermediate element may be input into the at least one lookup table for searching. The input of the at least one neural network module may be determined or obtained based on the search result and then input into the at least one neural network module for processing until all prediction processing steps are completed and the predicted pixel is obtained.
[0428] Optionally, the input to at least one neural network module can be determined based on the type, size parameter, and / or input range of at least one lookup table. For example, if two different types of lookup tables have different output ranges, their corresponding inputs to at least one neural network module will also be different; similarly, if two lookup tables with different size parameters have different output ranges, their corresponding inputs to at least one neural network module will also be different.
[0429] In this embodiment, by including at least one lookup branch and at least one neural network module in the lookup table structure, and the input of the at least one neural network module is determined or obtained according to the type, size parameters and / or input range of the at least one lookup table, it is possible to use at least one lookup branch of the lookup table and at least one neural network module to jointly perform prediction processing on at least one reference pixel or first reference intermediate element, and combine the advantages of both the lookup table in the lookup branch and the neural network in the neural network module to perform prediction processing, so as to improve the prediction effect.
[0430] Method 26: At least two lookup tables, where the output ranges of the two lookup tables are the same and / or the output ranges of the lookup tables are different;
[0431] Optionally, at least two lookup tables in the lookup table structure have the same output range.
[0432] Optionally, at least two lookup tables in the lookup table structure have different output ranges.
[0433] Optionally, the lookup table structure may contain at least three lookup tables, with two lookup tables having the same output range and two lookup tables having different output ranges.
[0434] For example, the lookup table structure includes lookup table 1, lookup table 2 and lookup table 3. Lookup table 1 and lookup table 2 have the same output range, while lookup table 3 has a different output range from lookup table 1 and lookup table 2.
[0435] Optionally, at least two lookup tables with different output ranges can be used to perform prediction processing on at least one reference pixel or the first reference intermediate element; or at least two lookup tables with the same output range can be used to perform prediction processing on at least one reference pixel or the first reference intermediate element.
[0436] In this embodiment, by including at least two lookup tables in the lookup table structure, and the output ranges of the at least two lookup tables being the same and / or different, it is possible to perform prediction processing more flexibly based on the output ranges of different lookup tables when performing prediction processing on at least one reference pixel or a first reference intermediate element according to the lookup table structure, thereby improving the prediction effect.
[0437] Method 27: At least two lookup tables, where the input ranges of the two lookup tables are the same and / or the input ranges of the lookup tables are different;
[0438] Optionally, at least two lookup tables in the lookup table structure have the same input range.
[0439] Optionally, at least two lookup tables in the lookup table structure have different input ranges.
[0440] Optionally, the lookup table structure may contain at least three lookup tables, with two lookup tables having the same input range and two lookup tables having different input ranges.
[0441] For example, the lookup table structure includes lookup table 1, lookup table 2 and lookup table 3. Lookup table 1 and lookup table 2 have the same input range, while lookup table 3 has a different input range than both lookup table 1 and lookup table 2.
[0442] Optionally, at least two lookup tables with different input ranges can be used to perform prediction processing on at least one reference pixel or a first reference intermediate element; and / or, at least two lookup tables with the same input range can be used to perform prediction processing on at least one reference pixel or a first reference intermediate element.
[0443] In this embodiment, by including at least two lookup tables in the lookup table structure, and having the same and / or different input ranges for the at least two lookup tables, it is possible to perform more flexible prediction processing based on the different input ranges of the lookup tables when performing prediction processing on at least one reference pixel or a first reference intermediate element according to the lookup table structure, thereby improving the prediction effect.
[0444] Method 28: At least one first residual module based on a lookup table, the first residual module including a first branch containing at least one lookup table structure, an adder, and a second branch containing a short-circuit structure;
[0445] Optionally, a residual structure can be set in the lookup table structure, that is, at least one first residual module based on the lookup table can be set, and the residual structure of the lookup table can be reflected through the first residual module.
[0446] Optionally, a first residual module can be used to process at least one reference pixel or the first reference intermediate element. For example, at least one reference pixel or the first reference intermediate element can be processed according to a first branch containing at least one lookup table structure and a second branch containing a short-circuit structure, and then the output results of the first branch and the second branch can be weighted and added together until all prediction processes are completed and the predicted pixel is obtained.
[0447] In this embodiment, by including at least one first residual module based on the lookup table structure, the first residual module includes a first branch containing at least one lookup table structure, an adder, and a second branch containing a short-circuit structure, it is possible to combine the residual structure with the prediction processing when performing prediction processing on at least one reference pixel or a first reference intermediate element based on the lookup table structure, so as to reflect the advantages of the residual structure and improve the prediction effect of the prediction processing.
[0448] Method 29: At least one second residual module based on a lookup table, the first input element of the second residual module passes through a first branch containing at least one lookup table structure to obtain a fourth output element, and the first input element of the third residual module passes through a second branch containing a short-circuit structure to obtain a fifth output element that is the same as the first input element. The fourth output element and the fifth output element are weighted and added by an adder to obtain a sixth output element.
[0449] Optionally, a residual structure can be set in the lookup table structure, such as at least one second residual module based on the lookup table, and / or at least one first residual module based on the lookup table.
[0450] Optionally, the first residual module can be the second residual module.
[0451] Optionally, the first input element (such as a pixel or index marker) of the second residual module can be determined or obtained based on at least one pixel to be predicted.
[0452] Optionally, in the first branch of the second residual module, the first input element can be input into at least one lookup table structure for processing, such as inputting its corresponding index into the lookup table for searching, and the search result is used as the fourth output element of the first branch.
[0453] Optionally, in the second branch of the second residual module, the first input element can be input into the second branch containing the short-circuit result, and the output can be the same as the first input element.
[0454] The fourth output element from the first branch and the fifth output element from the second branch are input into the adder for weighted summation to obtain the sixth output element.
[0455] Optionally, the sixth output element can be directly output as the predicted pixel, or the sixth output element can be used as an intermediate prediction element to continue subsequent prediction processing until all prediction processes are completed and the predicted pixel is obtained.
[0456] In this embodiment, by including at least one second residual module based on the lookup table structure, the first input element of the second residual module passes through a first branch containing at least one lookup table structure to obtain a fourth output element, and the first input element of the third residual module passes through a second branch containing a short-circuit structure to obtain a fifth output element that is the same as the first input element. The fourth output element and the fifth output element are weighted and added by an adder to obtain a sixth output element. This enables prediction processing to be combined with the residual structure when performing prediction processing on at least one reference pixel or a first reference intermediate element based on the lookup table structure, so as to reflect the advantages of the residual structure and improve the prediction effect of the prediction processing.
[0457] Sixth Embodiment
[0458] Based on any of the above embodiments, a sixth embodiment is proposed.
[0459] In this embodiment, the image processing method further includes: determining or obtaining the spatial features of at least one reference block after image preprocessing.
[0460] Optionally, at least one spatial feature can be determined or obtained based on at least one reference block after image preprocessing, and prediction processing can be performed on at least one image block based on the at least one spatial feature.
[0461] Optionally, this embodiment can be combined with any of the methods or steps in the above embodiments.
[0462] In this embodiment, by determining or obtaining at least one spatial feature based on at least one reference block after image preprocessing, and performing prediction processing on at least one image block based on at least one spatial feature, it is possible to comprehensively consider the spatial features of at least one reference block when performing prediction processing on at least one image block, so as to determine the key image block features that need to be processed in the reference block. In this way, the key image block features can be focused during prediction processing, which can improve the prediction effect of prediction processing on at least one image block.
[0463] Optionally, image preprocessing includes at least one of image sharpening, gradient calculation, transformation, and detection based on a detection operator.
[0464] Optionally, the spatial features of at least one reference block after image sharpening can be determined or obtained, and the at least one image block can be predicted based on the at least one spatial feature.
[0465] Optionally, the spatial features of at least one reference block after gradient calculation can be determined or obtained, and the at least one image block can be predicted based on the at least one spatial feature.
[0466] Optionally, the spatial features of at least one transformed reference block can be determined or obtained, and the at least one image block can be predicted based on the at least one spatial feature.
[0467] Alternatively, the transformation can be a Fourier transform, a wavelet transform, or a feature transformation network based on deep learning.
[0468] Alternatively, the transformation can also be a scale transformation, such as image scaling (bilinear interpolation, bicubic interpolation) or multi-scale pyramid construction (such as Gaussian pyramids, Laplacian pyramids).
[0469] Alternatively, the transformation can also be a linear transformation, a nonlinear transformation, or a projection transformation, such as the Hough transformation or perspective transformation.
[0470] Optionally, the spatial features of at least one reference block after texture detection based on the detection operator can be determined or obtained, and the at least one image block can be predicted based on the at least one spatial feature.
[0471] Optionally, the detection operator includes at least one of the Sobel operator, Scharr operator, Canny operator, and Laplacian operator.
[0472] Alternatively, the detection operator can also be a neural network-based operator.
[0473] The following example uses image sharpening to illustrate this.
[0474] Optionally, for video compression tasks, since more high-frequency information such as texture is lost, a 3x3 Unsharp Mask (USM) sharpening technique can be applied to image patches to enhance texture intensity as subsequent prior input. The USM-sharpened image is denoted as... right Spatial features are extracted at each position (i,j), such as Figure 7 As shown.
[0475] Optionally, for position (i,j), the surrounding 5×5 window is extracted. In order to make the most of the features of the entire window, the window is divided into inner and external, that is, the operation of the Mean module in the figure is performed. Then, the spatial features are obtained by averaging the two regions respectively.
[0476] Optionally, calculate the feature mean of position (i,j) and its upper, lower, left, and right positions, that is, calculate the feature mean of positions (i-1,j), (i+1,j), (i,j-1), and (i,j+1), denoted as a. inner Calculate the feature mean of the eight positions at the outermost layer of the window, that is, calculate the feature mean of positions (i-2,j-1), (i-2,j+1), (i-1,j-2), (i+1,j-2), (i+2,j-1), (i+2,j+1), (i-1,j+2), and (i+1,j+2), denoted as a. exter , (a inner a exter Let (i,j) be the spatial feature. Based on the spatial features of position (i,j), the spatial attention weights for position (i,j) can be further obtained:
[0477]
[0478] σ represents the sigmoid activation function; This can be either a Conv1x2 layer or a lookup table. Conv1x2 represents a 1x2 convolution kernel; style_feature is the spatial feature of the input, calculated by averaging the inner and external features. In other words, when obtaining the spatial attention weights at position (i,j) based on the spatial features, a convolutional layer (Conv1x2) or a lookup table is used, followed by a sigmoid layer. Since the output of the sigmoid function is between 0 and 1, it can be used as an attention weight / mask. The lookup table can be obtained by storing the input and output of the convolutional layer (Conv1x2).
[0479] Optionally, the result after spatial attention enhancement is:
[0480] f out =atten·f in +f in
[0481] f in It enhances the texture intensity as a priori input, and generates spatial attention weights through spatial feature extraction, further improving the network's ability to model spatial information and texture details.
[0482] In this embodiment, by performing image preprocessing based on at least one of the following: image sharpening, gradient calculation, transformation, and detection operators, and texture detection, and then performing prediction processing on at least one image block based on at least one spatial feature, it is possible to comprehensively consider the spatial features of at least one reference block when performing prediction processing on at least one image block, so as to determine the key image block features that need to be processed in the reference block. In this way, the prediction processing can focus on the key image block features, thereby improving the prediction effect of the prediction processing on at least one image block.
[0483] Fifth Embodiment
[0484] This application also provides a processing apparatus, referring to... Figure 8 The processing device includes:
[0485] The processing module A10 is used to perform prediction processing on at least one image block based on the spatial characteristics of at least one reference block.
[0486] Optionally, the spatial characteristics are determined or obtained based on at least one of the following:
[0487] Statistical feature values of a reference pixel and at least one neighboring pixel in at least one reference block;
[0488] Statistical feature values of at least two non-neighboring pixels in at least one reference block;
[0489] The statistical feature value of at least one of the following: a reference pixel in at least one reference block, at least one upper adjacent pixel, at least one left adjacent pixel, at least one right adjacent pixel, and at least one lower adjacent pixel of the reference pixel;
[0490] The reference pixel in at least one reference block, and the weighted statistical feature value of the pixels adjacent to the reference pixel;
[0491] The statistical feature value of at least one pixel of the outermost layer of the window centered on the reference pixel, where the window size parameter is less than or equal to the reference block size parameter;
[0492] The statistical feature values of eight pixels that are not adjacent to the reference pixel, and the eight pixels are not adjacent to each other;
[0493] The statistical feature value of a reference sub-block that contains at least one of the following: at least one reference pixel, at least one neighboring pixel, and at least one non-neighboring pixel;
[0494] At least one reference block contains statistical feature values of a reference region that include at least one reference pixel, at least one neighboring pixel, and at least one non-neighboring pixel.
[0495] Optionally, the reference block is determined or obtained based on at least one of the following:
[0496] At least one of the following: the top adjacent pixel, the top non-adjacent pixel, the left adjacent pixel, the left non-adjacent pixel, the top left adjacent pixel, and the top left non-adjacent pixel;
[0497] The image patch is corresponding to at least one of the following: neighboring block, non-neighboring block, cross-component block, co-position block, temporal block, and default block;
[0498] At least one of the following: width, height, block size, and block area of the image block;
[0499] Candidate motion vectors or candidate block vectors of image patches determine or generate candidate blocks;
[0500] If the first information of an image block satisfies the first condition, then the reference block is the first reference block;
[0501] If the first information of an image block satisfies the second condition, then the reference block is the second reference block.
[0502] Optionally, processing module A10 is configured to perform at least one of the following:
[0503] Based on at least one spatial feature and at least one lookup table, determine or obtain at least one spatial attention weight, and perform prediction processing on at least one image patch based on the at least one spatial attention weight.
[0504] Based on at least one spatial feature and at least one neural network, at least one spatial attention weight is determined or obtained, and at least one image patch is predicted based on the at least one spatial attention weight.
[0505] Based on at least one spatial feature and at least one activation function, at least one spatial attention weight is determined or obtained, and at least one image patch is predicted based on the at least one spatial attention weight.
[0506] Optionally, processing module A10 is configured to perform at least one of the following:
[0507] At least two first intermediate elements in the intermediate element block corresponding to at least one reference block have the same spatial attention weight and / or the spatial attention weights corresponding to the reference pixels are different.
[0508] The spatial attention weight corresponding to at least one first intermediate element in the intermediate element block corresponding to at least one reference block is greater than the spatial attention weight corresponding to another reference pixel;
[0509] The spatial attention weight corresponding to the intermediate element region of at least one reference block that contains at least one first intermediate element is greater than the spatial attention weight corresponding to the intermediate element region that contains at least one other first intermediate element.
[0510] The spatial attention weight corresponding to the intermediate element sub-block that contains at least one first intermediate element in the intermediate element block corresponding to at least one reference block is greater than the spatial attention weight corresponding to the intermediate element sub-block that contains at least one other first intermediate element.
[0511] Optionally, processing module A10 is configured to perform at least one of the following:
[0512] Based on at least one neural network, a lookup table structure, at least one item in the lookup table, and at least one reference block, determine or obtain at least one intermediate element block;
[0513] The lookup table structure includes at least one of the following:
[0514] At least one lookup table;
[0515] At least one neural network module;
[0516] There are at least two search branches, and at least one search branch has the same input as the other search branch;
[0517] There are at least two search branches, and the input of at least one search branch is determined or obtained based on the output of the other search branch;
[0518] At least two lookup branches, and at least two lookup branches are located in at least one channel corresponding to at least one reference pixel or the first reference intermediate element;
[0519] There are at least two search branches, and the search method for the search table with at least two search branches is parallel search and / or serial search;
[0520] At least two lookup tables, and the input of at least one lookup table is determined or obtained based on the type, size parameters and / or input range of the other lookup table;
[0521] At least two lookup tables, and at least two lookup tables are located in at least one channel corresponding to at least one reference pixel or the first reference intermediate element;
[0522] At least two lookup tables, and the lookup methods for at least two lookup tables within the same channel are parallel lookup and / or serial lookup;
[0523] At least two lookup tables, and the lookup tables in at least two channels are searched in parallel and / or serial manner;
[0524] At least one search branch and at least one neural network module, wherein the input of the at least one neural network module is the same as the input of the at least one search branch;
[0525] At least one search branch and at least one neural network module, wherein the input of the at least one neural network module is determined or obtained based on the output of the at least one search branch;
[0526] At least one lookup table and at least one neural network module, wherein the input of the at least one neural network module is determined or obtained based on the type, size parameters and / or input range of the at least one lookup table;
[0527] At least two lookup tables, at least two lookup tables have the same output range and / or the lookup tables have different output ranges;
[0528] At least two lookup tables, at least two lookup tables have the same input range and / or the lookup tables have different input ranges;
[0529] At least one first residual module based on a lookup table, the first residual module including a first branch containing at least one lookup table structure, an adder, and a second branch containing a short-circuit structure;
[0530] At least one second residual module based on a lookup table, the first input element of the second residual module is passed through a first branch containing at least one lookup table structure to obtain a fourth output element, and the first input element of the third residual module is passed through a second branch containing a short-circuit structure to obtain a fifth output element that is the same as the first input element. The fourth output element and the fifth output element are weighted and added by an adder to obtain a sixth output element.
[0531] Optionally, prediction processing is performed on at least one image patch based on at least one spatial attention weight, including at least one of the following:
[0532] Based on the spatial attention enhancement of the first intermediate element in at least one intermediate element block according to at least one spatial attention weight, at least one prediction block or prediction intermediate element is determined or obtained.
[0533] Based on at least one spatial attention weight, spatial attention enhancement is performed on at least one intermediate element region in at least one intermediate element block to determine or obtain at least one prediction block or prediction intermediate element.
[0534] Based on at least one spatial attention weight, the spatial attention enhancement is performed on at least one intermediate element sub-block in at least one intermediate element block to determine or obtain at least one prediction block or prediction intermediate element.
[0535] Optionally, the processing module A10 is used to: determine or obtain the spatial features of at least one image block after image preprocessing.
[0536] Optionally, image preprocessing includes at least one of the following:
[0537] Image sharpening;
[0538] Gradient calculation;
[0539] Transformation;
[0540] Detection is performed based on the detection operator.
[0541] The processing device provided in this application embodiment is similar in implementation principle and beneficial effect to the technical solution shown in the corresponding method embodiment above, and will not be described again here.
[0542] This application also provides a processing device, including a memory and a processor. The memory stores an image processing program, and when the image processing program is executed by the processor, it implements the steps of the image processing method in any of the above embodiments.
[0543] This application also provides a storage medium storing an image processing program, which, when executed by a processor, implements the steps of the image processing method in any of the above embodiments.
[0544] In the embodiments of the processing device and storage medium provided in this application, all the technical features of any of the above-described image processing method embodiments may be included. The extended and explanatory content of the specification is basically the same as that of the embodiments of the above methods, and will not be repeated here.
[0545] This application also provides a computer program product, which includes computer program code. When the computer program code is run on a computer, it causes the computer to perform the methods described in the various possible implementations above.
[0546] This application also provides a chip, including a memory and a processor. The memory is used to store a computer program, and the processor is used to call and run the computer program from the memory, so that a device with the chip installed performs the methods described in the various possible implementations above.
[0547] It is understood that the above scenarios are merely examples and do not constitute a limitation on the application scenarios of the technical solutions provided in the embodiments of this application. The technical solutions of this application can also be applied to other scenarios. For example, as those skilled in the art will know, with the evolution of system architecture and the emergence of new business scenarios, the technical solutions provided in the embodiments of this application are also applicable to similar technical problems.
[0548] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.
[0549] The steps in the method of this application embodiment can be adjusted, combined, or deleted according to actual needs.
[0550] The units in the device of this application embodiment can be merged, divided, and deleted according to actual needs.
[0551] In this application, the same or similar terms, concepts, technical solutions and / or application scenario descriptions are generally described in detail only when they appear for the first time. When they appear again, they are generally not repeated for the sake of brevity. When understanding the technical solutions and other contents of this application, the same or similar terms, concepts, technical solutions and / or application scenario descriptions that are not described in detail later can be referred to their previous relevant detailed descriptions.
[0552] In this application, the descriptions of the various embodiments have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0553] The technical features of the present application can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of the present application.
[0554] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) as described above, and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, controlled terminal, or network device, etc.) to execute the methods of each embodiment of this application.
[0555] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented, in whole or in part, as a computer program product. A computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the flow or function according to the embodiments of this application is generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a storage medium or transmitted from one storage medium to another. For example, computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) means. The storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, storage disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., a solid-state disk (SSD)).
[0556] The above are merely preferred embodiments of this application and do not limit the patent scope of this application. Any equivalent structural or procedural transformations made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of this application.
Claims
1. An image processing method, characterized in that, Including the following steps: S1, perform prediction processing on at least one image block based on the spatial characteristics of at least one reference block; Step S1 includes at least one of the following: Based on at least one spatial feature and at least one lookup table, determine or obtain at least one spatial attention weight, and perform prediction processing on at least one image patch based on the at least one spatial attention weight. Based on at least one spatial feature and at least one neural network, at least one spatial attention weight is determined or obtained, and at least one image patch is predicted based on the at least one spatial attention weight. Based on at least one spatial feature and at least one activation function, at least one spatial attention weight is determined or obtained, and at least one image patch is predicted based on the at least one spatial attention weight. The spatial attention weight is a weight coefficient that quantifies the importance of different spatial locations in the feature map of an image block. The prediction process for at least one image patch based on at least one spatial attention weight includes: Spatial attention enhancement is performed on the first intermediate element corresponding to at least one image block based on at least one spatial attention weight. Prediction is then performed based on the result of the spatial attention enhancement to obtain a prediction block. The first intermediate element is an intermediate element, which is a reference pixel or a pixel obtained by filtering, transforming, or downsampling the reference pixel.
2. The image processing method as described in claim 1, characterized in that, Spatial characteristics are determined or obtained based on at least one of the following: Statistical feature values of a reference pixel and at least one neighboring pixel in at least one reference block; Statistical feature values of at least two non-neighboring pixels in at least one reference block; The statistical feature value of at least one of the following: a reference pixel in at least one reference block, at least one upper adjacent pixel, at least one left adjacent pixel, at least one right adjacent pixel, and at least one lower adjacent pixel of the reference pixel; The reference pixel in at least one reference block, and the weighted statistical feature value of the pixels adjacent to the reference pixel; The statistical feature value of at least one pixel of the outermost layer of the window centered on the reference pixel, where the window size parameter is less than or equal to the reference block size parameter; The statistical feature values of eight pixels that are not adjacent to the reference pixel, and the eight pixels are not adjacent to each other; The statistical feature value of a reference sub-block that contains at least one of the following: at least one reference pixel, at least one neighboring pixel, and at least one non-neighboring pixel; At least one reference block contains statistical feature values of a reference region that include at least one reference pixel, at least one neighboring pixel, and at least one non-neighboring pixel.
3. The image processing method as described in claim 1, characterized in that, The reference block is determined or obtained based on at least one of the following: At least one of the following: the top adjacent pixel, the top non-adjacent pixel, the left adjacent pixel, the left non-adjacent pixel, the top left adjacent pixel, and the top left non-adjacent pixel; The image patch is corresponding to at least one of the following: neighboring block, non-neighboring block, cross-component block, co-position block, temporal block, and default block; At least one of the following: width, height, block size, and block area of the image block; Candidate motion vectors or candidate block vectors of image patches determine or generate candidate blocks; If the first information of an image block satisfies the first condition, then the reference block is the first reference block; If the first information of an image block satisfies the second condition, then the reference block is the second reference block.
4. The image processing method as described in claim 1, characterized in that, It also includes at least one of the following: At least two first intermediate elements in the intermediate element block corresponding to at least one reference block have the same and / or different spatial attention weights; The spatial attention weight corresponding to at least one first intermediate element in the intermediate element block corresponding to at least one reference block is greater than the spatial attention weight corresponding to another reference pixel; The spatial attention weight corresponding to the intermediate element region of at least one reference block that contains at least one first intermediate element is greater than the spatial attention weight corresponding to the intermediate element region that contains at least one other first intermediate element. The spatial attention weight corresponding to the intermediate element sub-block that contains at least one first intermediate element in the intermediate element block corresponding to at least one reference block is greater than the spatial attention weight corresponding to the intermediate element sub-block that contains at least one other first intermediate element. In this process, at least one reference pixel in at least one reference block is pre-filtered, transformed, or downsampled to obtain an intermediate element block.
5. The image processing method as described in claim 4, characterized in that, It also includes at least one of the following: Based on at least one neural network, a lookup table structure, at least one item in the lookup table, and at least one reference block, determine or obtain at least one intermediate element block; The lookup table structure includes at least one of the following: At least one lookup table; At least one neural network module; There are at least two search branches, and at least one search branch has the same input as the other search branch; There are at least two search branches, and the input of at least one search branch is determined or obtained based on the output of the other search branch; At least two lookup branches, and at least two lookup branches are located in at least one channel corresponding to at least one reference pixel or the first reference intermediate element; There are at least two search branches, and the search method for the search table with at least two search branches is either parallel search or serial search; At least two lookup tables, and the input of at least one lookup table is determined or obtained based on the type, size parameters and / or input range of the other lookup table; At least two lookup tables, and at least two lookup tables are located in at least one channel corresponding to at least one reference pixel or the first reference intermediate element; At least two lookup tables are required, and the lookup methods for the at least two lookup tables within the same channel are either parallel lookup or serial lookup; At least two lookup tables, and the lookup tables in at least two channels are searched in parallel and / or serial manner; At least one search branch and at least one neural network module, wherein the input of the at least one neural network module is the same as the input of the at least one search branch; At least one search branch and at least one neural network module, wherein the input of the at least one neural network module is determined or obtained based on the output of the at least one search branch; At least one lookup table and at least one neural network module, wherein the input of the at least one neural network module is determined or obtained based on the type, size parameters and / or input range of the at least one lookup table; At least two lookup tables, at least two lookup tables have the same output range and / or the lookup tables have different output ranges; At least two lookup tables, at least two lookup tables have the same input range and / or the lookup tables have different input ranges; At least one first residual module based on a lookup table, the first residual module including a first branch containing at least one lookup table structure, an adder, and a second branch containing a short-circuit structure; At least one second residual module based on a lookup table, the first input element of the second residual module is passed through a first branch containing at least one lookup table structure to obtain a fourth output element, and the first input element of the third residual module is passed through a second branch containing a short-circuit structure to obtain a fifth output element that is the same as the first input element. The fourth output element and the fifth output element are weighted and added by an adder to obtain a sixth output element.
6. The image processing method as described in claim 4, characterized in that, Prediction processing of at least one image patch based on at least one spatial attention weight, further comprising at least one of the following: Based on at least one spatial attention weight, at least one prediction block or prediction intermediate element is determined or obtained by performing spatial attention enhancement on at least one intermediate element region in at least one intermediate element block; wherein, prediction is performed based on the result of spatial attention enhancement on at least one intermediate element region in at least one intermediate element to determine or obtain at least one prediction block or prediction intermediate element. Based on at least one spatial attention weight, at least one prediction block or prediction intermediate element is determined or obtained by performing spatial attention enhancement on at least one intermediate element sub-block in at least one intermediate element block; wherein, prediction is performed based on the result of spatial attention enhancement on at least one intermediate element sub-block in at least one intermediate element block to determine or obtain at least one prediction block or prediction intermediate element. The intermediate prediction element is the intermediate output generated during the prediction process.
7. The image processing method according to any one of claims 1 to 3, characterized in that, Also includes: Determine or obtain the spatial features of at least one reference block after image preprocessing.
8. The image processing method as described in claim 7, characterized in that, Image preprocessing includes at least one of the following: Image sharpening; Gradient calculation; Transformation; Detection is performed based on the detection operator.
9. A processing device, characterized in that, include: A memory and a processor, wherein the memory stores an image processing program, and the image processing program, when executed by the processor, implements the steps of the image processing method as described in any one of claims 1 to 8.
10. A storage medium, characterized in that, The storage medium stores a computer program, which, when executed by a processor, implements the steps of the image processing method as described in any one of claims 1 to 8.