Processing method, processing device, and storage medium
By dividing the blocks to be predicted in video encoding into multiple prediction areas and selecting appropriate prediction models and parameters for each area, the problem of low prediction accuracy in the prior art is solved, and the efficiency of video encoding and decoding is improved.
Patent Information
- Application Number
- PCT/CN2023/133391
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-11-22
- Publication Date
- 2025-05-30
AI Technical Summary
During block prediction processing, the existing video encoding standards use the same prediction model and model parameters for all pixels of the block through cross-component prediction technology, and fail to consider the different characteristics of different pixels in the block, resulting in a decrease in prediction accuracy, which in turn limits the efficiency of video encoding and decoding.
A processing method is proposed by dividing the block to be predicted into multiple prediction areas and making predictions based on the prediction model and model parameters of each prediction area. This method avoids the limitation of each predicted region size by power of 2 and allows different prediction models and parameters to be adopted in different prediction regions.
The prediction accuracy of block prediction is improved, thereby improving the efficiency of video encoding and decoding.
Smart Images

Figure CN2023133391_30052025_PF_FP_ABST
Abstract
Description
Processing method, processing device and storage medium Technical Field
[0001] The present application relates to the field of image processing technology, and in particular to a processing method, a processing device, and a storage medium. Background Art
[0002] The existing High-Efficiency Video Coding standard protocol (H.266 / VVC) proposes a video frame encoding technology to improve encoding performance without significantly increasing computational complexity. Specifically, when encoding and decoding video frames, the protocol divides each frame into different blocks, performs prediction processing, and then performs encoding and decoding processing.
[0003] During the process of conceiving and implementing this application, the inventors discovered that there are at least the following problems: when performing block prediction processing, only the same prediction model and the same set of model parameters can be used for all pixels of the block through cross-component prediction technology, without taking into account that different pixels in the block have different characteristics, resulting in reduced prediction accuracy, which in turn limits the efficiency of video encoding and decoding.
[0004] The preceding description is intended to provide general background information and does not necessarily constitute prior art.
[0005] Summary of the Invention
[0006] In response to the above technical problems, the present application provides a processing method, a processing device and a storage medium, aiming to solve the technical problem of how to improve the prediction accuracy of block prediction and thereby improve the efficiency of video encoding and / or decoding.
[0007] This application provides a processing method that can be applied to a processing device (such as an intelligent terminal or a server, etc.), including the following steps:
[0008] S10, dividing the block to be predicted into at least one prediction area;
[0009] S20: performing prediction based on the prediction model and / or model parameters corresponding to the prediction area.
[0010] Optionally, the prediction model comprises a cross-component prediction model.
[0011] Optionally, step S10 includes at least one of the following:
[0012] Dividing the block to be predicted into at least one prediction area according to size information of the block to be predicted;
[0013] Dividing the block to be predicted into at least one prediction area according to neighboring pixels of the block to be predicted;
[0014] Dividing the block to be predicted into at least one prediction area according to the division information of the block to be predicted;
[0015] Dividing the second component and / or the third component of the block to be predicted into at least one prediction area;
[0016] The block to be predicted is divided into at least one prediction area according to pixel gradient values of neighboring blocks.
[0017] Optionally, the step S20 includes the following steps:
[0018] S21, performing cross-component prediction on at least one prediction area according to the prediction model and / or model parameters corresponding to the at least one prediction area, and determining or generating a prediction result;
[0019] S22: Determine or generate a prediction block according to the prediction results of the plurality of prediction areas.
[0020] Optionally, step S22 includes at least one of the following:
[0021] Determine or obtain a second component prediction block according to prediction results of the plurality of second component prediction areas;
[0022] A third component prediction block is determined or obtained according to the prediction results of the plurality of third component prediction areas.
[0023] Optionally, the processing method further includes at least one of the following:
[0024] Reconstructing the block based on the prediction residual obtained from the bitstream and the second component prediction block, and determining or generating a reconstruction of the block;
[0025] Reconstructing the block based on the prediction residual obtained from the bitstream and the third component prediction block, and determining or generating a reconstruction of the block;
[0026] Determining or generating a prediction residual based on the pixels of the block to be predicted and the second component prediction block;
[0027] A prediction residual is determined or generated according to the pixels of the block to be predicted and the third component prediction block.
[0028] Optionally, the prediction region includes a second component prediction region and / or a third component prediction region.
[0029] Optionally, the processing method further includes at least one of the following:
[0030] The second component prediction area includes a rectangular area and / or a non-rectangular area;
[0031] The third component prediction area includes a rectangular area and / or a non-rectangular area;
[0032] Determining or obtaining a prediction model and / or model parameters corresponding to at least one prediction area based on reference pixels of the block to be predicted;
[0033] Determining or obtaining a cross-component prediction model and / or model parameters of the second component prediction region based on reference pixels of the second component prediction region and / or reference pixels of the third component prediction region;
[0034] The cross-component prediction model and / or model parameters of the third component prediction region are determined or obtained according to the reference pixels of the second component prediction region and / or the reference pixels of the third component prediction region.
[0035] Optionally, the processing method further includes at least one of the following:
[0036] The reference pixels of the second component prediction area include neighboring pixels of the first component of the block to be predicted and neighboring pixels of the second component of the block to be predicted;
[0037] The reference pixels of the third component prediction area include neighboring pixels of the first component of the block to be predicted and neighboring pixels of the third component of the block to be predicted.
[0038] The present application also provides a processing device, comprising:
[0039] A division module, configured to divide the block to be predicted into at least one prediction area;
[0040] The prediction module is used to make predictions based on the prediction model and / or model parameters corresponding to the prediction area.
[0041] The present application also provides a processing device, including: a memory and a processor, wherein a processing program is stored in the memory, and when the processing program is executed by the processor, the steps of any of the above processing methods are implemented.
[0042] The processing device mentioned in this application can be a smart terminal or a server, etc.
[0043] The present application also provides a storage medium, wherein the storage medium stores a computer program, and when the computer program is executed by a processor, the steps of any of the processing methods described above are implemented.
[0044] As described above, the processing method of the present application can be applied to a processing device, including: dividing the block to be predicted into at least one prediction area; and making predictions based on the prediction model and / or model parameters corresponding to at least one prediction area. Through the technical solution of the present application, it is possible to divide the block to be predicted into multiple prediction areas during the prediction stage, thereby avoiding the phenomenon that the size of each prediction area is limited by a power of 2, and / or the prediction area is rectangular, and the prediction is made based on the prediction model and / or model parameters corresponding to the prediction area, such as making cross-component predictions based on a cross-component prediction model, thereby avoiding the phenomenon that all pixels of the block are predicted using a fixed cross-component prediction model and model parameters, resulting in low prediction accuracy, thereby improving the prediction accuracy of block prediction, and thereby improving the efficiency of video encoding and / or decoding. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] The accompanying drawings herein are incorporated into and constitute a part of the specification, illustrate embodiments consistent with the present application, and together with the specification, are used to explain the principles of the present application. In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following is a brief introduction to the drawings required for describing the embodiments. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without inventive work.
[0046] FIG1 is a schematic diagram of the hardware structure of a mobile terminal for implementing various embodiments of the present application;
[0047] FIG2 is a diagram of a communication network system architecture provided by an embodiment of the present application;
[0048] FIG3 is a schematic flow chart of a processing method according to the first embodiment;
[0049] FIG4 is a schematic diagram of prediction area division according to the processing method of the first embodiment;
[0050] FIG5 is a schematic diagram of a scene of a current frame in the processing method according to the first embodiment;
[0051] FIG6 is a schematic diagram of a scene of the Y component of the current frame in the processing method according to the first embodiment;
[0052] FIG7 is a schematic diagram of a scene of a U component of a current frame in a processing method according to the first embodiment;
[0053] FIG8 is a schematic diagram of a scene of a V component of a current frame in a processing method according to the first embodiment;
[0054] FIG9 is a schematic diagram of dividing the prediction area according to the neighbor gradient values in the processing method according to the first embodiment;
[0055] FIG10 is a schematic flow chart of a processing method according to a second embodiment;
[0056] 11 is a schematic diagram of cross-component prediction in multiple prediction regions at the decoding side according to a processing method according to a fourth embodiment;
[0057] 12 is a schematic diagram illustrating cross-component prediction in multiple prediction regions at the encoding side according to a processing method according to a fourth embodiment;
[0058] FIG13 is a schematic diagram of a prediction region based on CCCM according to a processing method shown in the fifth embodiment;
[0059] 14 is a schematic diagram of components to be predicted and reference pixels of a current block according to a processing method shown in the fifth embodiment;
[0060] 15 is another schematic diagram of a component to be predicted and its reference pixels of a current block according to a processing method according to the fifth embodiment;
[0061] FIG16 is a schematic diagram of modules of a processing device according to an embodiment of the present application.
[0062] The purpose of this application, its features, and advantages will be further described in conjunction with the embodiments and with reference to the accompanying drawings. The accompanying drawings illustrate specific embodiments of this application, which will be described in more detail below. These drawings and the accompanying text are not intended to limit the scope of the present application in any way, but rather to illustrate the concepts of this application to those skilled in the art by reference to specific embodiments. DETAILED DESCRIPTION
[0063] Exemplary embodiments will be described in detail herein, with examples illustrated in the accompanying drawings. In the following description, when referring to the drawings, identical numerals in different figures represent identical or similar elements, unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all embodiments consistent with the present application. Rather, they are merely examples of apparatus and methods consistent with certain aspects of the present application, as detailed in the appended claims.
[0064] It should be noted that, in this document, the terms "comprises", "includes" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the sentence "comprises a ..." does not exclude the presence of other identical elements in the process, method, article or device comprising the element, and / or, components, features, and elements with the same name in different embodiments of the present application may have the same meaning or different meanings, and their specific meanings need to be determined by their explanation in the specific embodiment or further combined with the context of the specific embodiment.
[0065] It should be understood that although the terms first, second, third, etc. may be used herein to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of this document, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the term "if" as used herein may be interpreted as "when," "when," or "in response to a determination." Furthermore, as used herein, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context indicates otherwise. It should be further understood that the terms "comprising" and "including" indicate the presence of the described features, steps, operations, elements, components, items, types, and / or groups, but do not exclude the presence, occurrence, or addition of one or more other features, steps, operations, elements, components, items, types, and / or groups. The terms "or," "and / or," "including at least one of the following," etc., used herein, may be interpreted as inclusive, or mean any one or any combination. For example, “comprising at least one of the following: A, B, C” means “any of the following: A; B; C; A and B; A and C; B and C; A and B and C”; and for another example, “A, B or C” or “A, B and / or C” means “any of the following: A; B; C; A and B; A and C; B and C; A and B and C”. An exception to this definition will occur only when a combination of elements, functions, steps or operations are inherently mutually exclusive in some manner.
[0066] It should be understood that, although the various steps in the flowchart in the embodiment of the present application are shown in sequence according to the indication of the arrows, these steps are not necessarily performed in sequence in the order indicated by the arrows. Unless clearly stated herein, the execution of these steps is not strictly limited in order, and they can be performed in other orders. Moreover, at least a portion of the steps in the figure may include multiple sub-steps or multiple stages, and these sub-steps or stages are not necessarily performed at the same time, but can be performed at different times, and their execution order is not necessarily performed in sequence, but can be performed in turn or alternately with at least a portion of other steps or sub-steps or stages of other steps.
[0067] As used herein, the words "if" and "if" may be interpreted as "at the time of" or "when" or "in response to determining" or "in response to detecting," depending on the context. Similarly, the phrases "if it is determined" or "if (stated condition or event) is detected" may be interpreted as "when it is determined" or "in response to the determination" or "when detecting (stated condition or event)" or "in response to detecting (stated condition or event)," depending on the context.
[0068] It should be noted that in this article, step codes such as S10 and S20 are used for the purpose of expressing the corresponding content more clearly and concisely, and do not constitute a substantial limitation on the order. When implementing the step, those skilled in the art may execute S20 first and then S10, etc., but these should all be within the scope of protection of this application.
[0069] It should be understood that the specific embodiments described herein are only used to explain the present application and are not intended to limit the present application.
[0070] In the subsequent description, the use of suffixes such as "module", "component" or "unit" to represent elements is only for the purpose of facilitating the description of the present application and has no specific meaning. Therefore, "module", "component" or "unit" can be used interchangeably.
[0071] The processing devices mentioned in this application may be smart terminals or servers. Smart terminals may be implemented in various forms. For example, the smart terminals described in this application may include smart terminals such as mobile phones, tablet computers, laptop computers, PDAs, portable media players (PMPs), navigation devices, wearable devices, smart bracelets, pedometers, etc., as well as fixed terminals such as digital TVs and desktop computers.
[0072] The subsequent description will be made by taking a mobile terminal as an example. It will be understood by those skilled in the art that, in addition to components specifically used for mobile purposes, the configuration according to the embodiments of the present application can also be applied to fixed-type terminals.
[0073] Please refer to Figure 1, which is a schematic diagram of the hardware structure of a mobile terminal for implementing various embodiments of the present application. The mobile terminal 100 may include components such as an RF (Radio Frequency) unit 101, a WiFi module 102, an audio output unit 103, an A / V (Audio / Video) input unit 104, a sensor 105, a display unit 106, a user input unit 107, an interface unit 108, a memory 109, a processor 110, and a power supply 111. Those skilled in the art will understand that the mobile terminal structure shown in Figure 1 does not limit the mobile terminal. The mobile terminal may include more or fewer components than shown, or may combine certain components, or arrange the components differently.
[0074] The following is a detailed introduction to the various components of the mobile terminal in conjunction with Figure 1:
[0075] The RF unit 101 can be used to send and receive information or receive signals during calls. Specifically, it receives downlink information from the base station and transmits it to the processor 110 for processing. It also transmits uplink data to the base station. Typically, the RF unit 101 includes, but is not limited to, an antenna, at least one amplifier, a transceiver, a coupler, a low-noise amplifier, a duplexer, and / or other components. Furthermore, the RF unit 101 can communicate with the network and other devices via wireless communication. The above-mentioned wireless communications can use any communication standard or protocol, including but not limited to GSM (Global System of Mobile communication), GPRS (General Packet Radio Service), CDMA2000 (Code Division Multiple Access 2000), WCDMA (Wideband Code Division Multiple Access), TD-SCDMA (Time Division-Synchronous Code Division Multiple Access), FDD-LTE (Frequency Division Duplexing-Long Term Evolution), TDD-LTE (Time Division Duplexing-Long Term Evolution), 5G and 6G, etc.
[0076] WiFi is a short-range wireless transmission technology. A mobile terminal, through WiFi module 102, enables users to send and receive emails, browse web pages, and access streaming media, providing wireless broadband Internet access. Although FIG1 illustrates WiFi module 102, it is understood that it is not a required component of the mobile terminal and can be omitted as needed without altering the essence of the invention.
[0077] The audio output unit 103 can convert audio data received by the RF unit 101 or the WiFi module 102 or stored in the memory 109 into an audio signal and output it as sound when the mobile terminal 100 is in a call signal reception mode, a talk mode, a recording mode, a voice recognition mode, a broadcast reception mode, or the like. Furthermore, the audio output unit 103 can also provide audio output related to a specific function performed by the mobile terminal 100 (e.g., a call signal reception sound, a message reception sound, etc.). The audio output unit 103 may include a speaker, a buzzer, or the like.
[0078] The A / V input unit 104 is used to receive audio or video signals. The A / V input unit 104 may include a graphics processing unit (GPU) 1041 and a microphone 1042. The GPU 1041 processes image data of still images or videos captured by an image capture device (e.g., a camera) in video capture mode or image capture mode. The processed image frames may be displayed on the display unit 106. The image frames processed by the GPU 1041 may be stored in the memory 109 (or other storage medium) or transmitted via the RF unit 101 or the WiFi module 102. The microphone 1042 may receive sound (audio data) in operating modes such as phone call mode, recording mode, and voice recognition mode, and may process such sound into audio data. In phone call mode, the processed audio (voice) data may be converted into a format that can be transmitted to a mobile communication base station via the RF unit 101. The microphone 1042 may implement various types of noise cancellation (or suppression) algorithms to eliminate (or suppress) noise or interference generated during the reception and transmission of audio signals.
[0079] The mobile terminal 100 also includes at least one sensor 105, such as a light sensor, a motion sensor, and other sensors. Optionally, the light sensor includes an ambient light sensor and a proximity sensor. Optionally, the ambient light sensor can adjust the brightness of the display panel 1061 according to the brightness of the ambient light, and the proximity sensor can turn off the display panel 1061 and / or the backlight when the mobile terminal 100 is moved to the ear. As a type of motion sensor, the accelerometer sensor can detect the magnitude of acceleration in all directions (generally three axes), and can detect the magnitude and direction of gravity when stationary. It can be used for applications that recognize the posture of the mobile phone (such as horizontal and vertical screen switching, related games, magnetometer posture calibration), vibration recognition related functions (such as pedometer, tapping), etc.; as for other sensors that can be configured in the mobile phone, such as fingerprint sensors, pressure sensors, iris sensors, molecular sensors, gyroscopes, barometers, hygrometers, thermometers, infrared sensors, etc., they will not be described here.
[0080] The display unit 106 is used to display information input by the user or information provided to the user. The display unit 106 may include a display panel 1061, which may be configured in the form of a liquid crystal display (LCD), an organic light-emitting diode (OLED), or the like.
[0081] The user input unit 107 can be used to receive input digital or character information and generate key signal input related to user settings and function control of the mobile terminal. Optionally, the user input unit 107 may include a touch panel 1071 and other input devices 1072. The touch panel 1071, also known as a touch screen, can collect user touch operations on or near it (such as operations performed by the user using a finger, stylus, or any other suitable object or accessory on or near the touch panel 1071) and drive corresponding connected devices according to a pre-set program. The touch panel 1071 may include two parts: a touch detection device and a touch controller. Optionally, the touch detection device detects the user's touch direction and detects the signal generated by the touch operation, and transmits the signal to the touch controller; the touch controller receives the touch information from the touch detection device and converts it into touch point coordinates, which are then sent to the processor 110. It can also receive and execute commands sent by the processor 110. And / or, the touch panel 1071 can be implemented using various types such as resistive, capacitive, infrared, and surface acoustic wave. In addition to the touch panel 1071, the user input unit 107 may further include other input devices 1072. Optionally, the other input devices 1072 may include, but are not limited to, one or more of a physical keyboard, function keys (such as volume control keys, power keys, etc.), a trackball, a mouse, a joystick, etc., and the specifics are not limited here.
[0082] Optionally, the touch panel 1071 may overlay the display panel 1061. When the touch panel 1071 detects a touch operation on or near it, it transmits the information to the processor 110 to determine the type of touch event. The processor 110 then provides a corresponding visual output on the display panel 1061 based on the type of touch event. Although in FIG1 , the touch panel 1071 and the display panel 1061 are shown as two separate components to implement the input and output functions of the mobile terminal, in some embodiments, the touch panel 1071 and the display panel 1061 may be integrated to implement the input and output functions of the mobile terminal, which is not limited to this specific embodiment.
[0083] The interface unit 108 serves as an interface through which at least one external device can be connected to the mobile terminal 100. For example, the external device may include a wired or wireless headset port, an external power supply (or battery charger) port, a wired or wireless data port, a memory card port, a port for connecting a device with an identification module, an audio input / output (I / O) port, a video I / O port, a headphone port, etc. The interface unit 108 may be used to receive input (e.g., data information, power, etc.) from an external device and transmit the received input to one or more elements within the mobile terminal 100 or may be used to transmit data between the mobile terminal 100 and an external device.
[0084] Memory 109 can be used to store software programs and various data. Memory 109 may primarily include a program storage area and a data storage area. Optionally, the program storage area may store an operating system and at least one application required for a function (such as a sound playback function or an image playback function); the data storage area may store data generated based on the use of the mobile phone (such as audio data, a phone book, etc.). Furthermore, / or, memory 109 may include high-speed random access memory and non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device.
[0085] Processor 110 is the control center of the mobile terminal, connecting all components of the mobile terminal using various interfaces and circuits. By running or executing software programs and / or modules stored in memory 109 and accessing data stored in memory 109, it executes various functions of the mobile terminal and processes data, thereby providing overall monitoring of the mobile terminal. Processor 110 may include one or more processing units; preferably, processor 110 may integrate an application processor and a modem processor. Optionally, the application processor primarily handles the operating system, user interface, and application programs, while the modem processor primarily handles wireless communications. It is understood that the modem processor may not be integrated into processor 110.
[0086] The mobile terminal 100 may also include a power supply 111 (such as a battery) for supplying power to various components. Preferably, the power supply 111 may be logically connected to the processor 110 through a power management system, thereby managing functions such as charging, discharging, and power consumption through the power management system.
[0087] Although not shown in FIG. 1 , the mobile terminal 100 may further include a Bluetooth module, etc., which will not be described in detail here.
[0088] To facilitate understanding of the embodiments of the present application, the communication network system on which the mobile terminal of the present application is based is described below.
[0089] Please refer to Figure 2, which is a communication network system architecture diagram provided in an embodiment of the present application. The communication network system is an LTE system of universal mobile communication technology. The LTE system includes a UE (User Equipment) 201, an E-UTRAN (Evolved UMTS Terrestrial Radio Access Network) 202, an EPC (Evolved Packet Core) 203 and an operator's IP service 204, which are connected in sequence.
[0090] Optionally, UE201 may be the above-mentioned terminal 100, which will not be described in detail here.
[0091] E-UTRAN 202 includes eNodeB 2021 and other eNodeBs 2022 . Optionally, eNodeB 2021 may be connected to other eNodeBs 2022 via a backhaul (eg, an X2 interface). eNodeB 2021 is connected to EPC 203 , and eNodeB 2021 may provide access from UE 201 to EPC 203 .
[0092] EPC 203 may include an MME (Mobility Management Entity) 2031, an HSS (Home Subscriber Server) 2032, other MMEs 2033, an SGW (Serving Gate Way) 2034, a PGW (PDN Gate Way) 2035, and a PCRF (Policy and Charging Rules Function) 2036. Optionally, MME 2031 is a control node that processes signaling between UE 201 and EPC 203, providing bearer and connection management. HSS 2032 provides registers for managing functions such as the Home Location Register (not shown) and stores user-specific information such as service features and data rates. All user data can be sent through SGW2034, PGW2035 can provide IP address allocation and other functions for UE 201, PCRF2036 is the policy and charging control policy decision point for service data flow and IP bearer resources, and it selects and provides available policy and charging control decisions for the policy and charging execution function unit (not shown in the figure).
[0093] The IP service 204 may include the Internet, an intranet, an IMS (IP Multimedia Subsystem), or other IP services.
[0094] Although the above introduction takes the LTE system as an example, those skilled in the art should know that this application is not only applicable to the LTE system, but can also be applied to other wireless communication systems, such as GSM, CDMA2000, WCDMA, TD-SCDMA, 5G and future new network systems (such as 6G), etc., which are not limited here.
[0095] Based on the above-mentioned mobile terminal hardware structure and communication network system, various embodiments of the present application are proposed.
[0096] First embodiment
[0097] 3 , which is a flow chart of a processing method according to a first embodiment, the processing method of the embodiment of the present application can be applied to a processing device, including the following steps:
[0098] S10, dividing the block to be predicted into at least one prediction area;
[0099] S20: performing prediction based on the prediction model and / or model parameters corresponding to the prediction area.
[0100] In this embodiment, the processing device may first determine the coding block to be predicted. The block to be predicted may be a frame of image or a segmented image block of a frame of image, without limitation. Alternatively, the processing device may be a smart terminal, such as a mobile phone or computer, or a server or cloud server. Alternatively, the block to be predicted may be a block to be encoded on the encoding side or a block to be decoded on the decoding side.
[0101] Alternatively, because different pixels in each block have different characteristics, and in the H.266 / VCC protocol or video codec technology, cross-component prediction technology uses the same prediction model and the same set of model parameters for all pixels in each block, which leads to the following defects:
[0102] Defect 1: It is impossible to use different reference pixels for different pixels in each block.
[0103] Defect 2: The division of blocks imposes restrictions on the division rules.
[0104] For example, when using binary, ternary, and quadtree partitioning, the blocks must be rectangular, and both the width and height of the blocks must meet the idempotence requirement of 2. For example, if you want to use different prediction models and parameters for a 64x64 block, you need to partition the 64x64 block into four 32x32 blocks. Each 32x32 block can use a different prediction model and parameters, but the size of each 32x32 block must be a power of 2, and the transforms for each 32x32 block after partitioning are independent. Defects 1 and 2 hinder improvements in encoding and decoding efficiency.
[0105] Optionally, in view of the current situation where all pixels in a block use the same cross-component model and model parameters, in this embodiment, during the prediction process of the block, the block is divided into multiple prediction areas, such as cross-component prediction areas.
[0106] Optionally, since the division of the prediction area only occurs in the prediction stage, the size of the prediction area obtained by the division is not limited by the power of 2, and the divided blocks can be non-rectangular areas. Optionally, the non-rectangular area may include a triangular area, a non-rectangular quadrilateral area, a pentagonal area, or a non-rectangular polygonal area composed of at least two rectangular areas, etc. And / or, the non-rectangular area can also be an irregularly shaped area, which is not limited here. Optionally, different cross-component prediction areas have different reference areas, that is, different prediction models and / or model parameters can be used. In addition, it is possible to adopt a corresponding prediction method for each prediction area, such as cross-component prediction, thereby improving the prediction accuracy, and thereby improving the encoding efficiency and / or decoding efficiency.
[0107] Optionally, in current technology, a CTU (coding tree unit) can be divided into multiple CUs (coding units), also known as blocks. Each CU can be further divided into blocks to obtain at least two blocks, and each block is predicted and reconstructed separately. All reconstructions are then weighted summed and / or combined and / or spliced to obtain the final result. For example, block 1 after division is predicted to obtain prediction 1, and prediction 1 is reconstructed to obtain reconstruction 1. Block 2 after division is predicted to obtain prediction 2, and prediction 2 is reconstructed to obtain reconstruction 2. Reconstruction 1 and reconstruction 2 are weighted summed and / or combined and / or spliced to obtain the reconstruction of the CU. Optionally, the prediction area is divided into prediction areas by dividing the prediction block (e.g., a CU) into prediction areas to obtain at least two prediction areas. Prediction is performed on the two prediction areas, and the prediction results are weighted summed and / or combined and / or spliced before reconstruction. For example, prediction area 1 is predicted to obtain prediction 1. Prediction area 2 is predicted to obtain prediction 2. The prediction 1 and the prediction 2 are weightedly summed and / or combined and / or concatenated to obtain a block prediction, and the block prediction is reconstructed to obtain a block reconstruction.
[0108] Optionally, before dividing the blocks to be predicted, the processing device may pre-store various images and videos and select an image to be predicted from each image as the block to be predicted, or segment the selected image and use the segmented image blocks as coding blocks. Alternatively, a frame of image may be extracted from a video sequence as the block to be predicted, or the extracted frame of image may be segmented to obtain image blocks, which are then used as the blocks to be predicted when performing coding prediction on the image blocks. Alternatively, the processing device may receive an input image or video and extract a frame of image from the image or video as the block to be predicted, or segment the extracted frame of image to obtain image blocks as the blocks to be predicted. Alternatively, the processing device may receive an image or video sent from another network device and extract a frame of image from the image or video as the block to be predicted, or segment the extracted frame of image to obtain image blocks as the blocks to be predicted. In this case, the processing device may pre-establish a communication connection with a network device on the network side of the mobile communication system in which it is located. Thus, the network device can transmit the image or video to the terminal device via this communication connection, and the terminal device then receives the image or video. Optionally, the processing device obtains the block to be predicted from the bitstream.
[0109] Optionally, the block to be predicted is divided into at least one prediction region, and prediction is performed on the prediction region using a prediction model and / or model parameters corresponding to the prediction region, such as intra-frame prediction and / or inter-frame prediction. The final results are weighted summed and / or combined and / or concatenated to obtain a more accurate prediction result for the block. For example, a prediction block is obtained by prediction on the encoding side, and a reconstruction of the block is obtained by prediction on the decoding side.
[0110] Optionally, this embodiment can divide the block to be predicted into multiple prediction regions, and perform cross-component prediction, same-component prediction, intra-frame prediction, or inter-frame prediction on each prediction region. Optionally, prediction is performed on each prediction region, and the prediction results are concatenated to obtain a prediction result for the block to be predicted. When performing prediction on each prediction region, each prediction region can calculate its own model parameters based on its own reference pixels, and then perform prediction based on the model parameters.
[0111] Optionally, after acquiring the image block, the processing device may perform a division process on the image block in the division phase, and may perform a division process on the divided block again in the prediction phase to obtain at least one prediction region.
[0112] Optionally, after the processing device obtains the block to be predicted, the block to be predicted may be divided into at least one prediction region according to a pre-set block division mode.
[0113] Optionally, the block division mode may be a division rule set in advance, and the block to be predicted is divided into at least one prediction area according to the division rule. Optionally, the division rule may be a rule set in advance, such as dividing the block to be predicted according to a certain size ratio; such as determining or generating at least one dividing line within the block to be predicted, and dividing the block to be predicted into at least one prediction area according to the dividing line; such as selecting at least one position point within the block to be predicted as a dividing point, and extending at least two non-overlapping dividing lines from the dividing point to the edge of the block to be predicted, and dividing the block to be predicted into at least one prediction area according to the dividing line; or it may be a rule set based on other methods, which is not limited here. Optionally, the division rule may be an approximately uniform horizontal division, an approximately uniform vertical division, a diagonal division, an unequal division, etc.
[0114] Optionally, the prediction area may be a rectangular area and / or a non-rectangular area.
[0115] Optionally, the non-rectangular area may include a triangular area, a non-rectangular quadrilateral area, a pentagonal area, or a non-rectangular polygonal area composed of at least two rectangular areas, etc. And / or, the non-rectangular area may also be an irregularly shaped area, which is not limited here.
[0116] Optionally, when dividing the prediction block, the prediction block may be divided at least once.
[0117] Optionally, the block partitioning mode selected each time the partitioning is performed may be the same or different.
[0118] Optionally, when the block to be predicted is divided multiple times, a geometric division method combined with a prediction mode of the block may be selected for division until at least one prediction region is obtained.
[0119] Optionally, the geometric partitioning may include partitioning the block to be predicted by geometrically positioned partitioning lines.
[0120] Optionally, step S10 includes at least one of the following methods 1 to 5:
[0121] Method 1: dividing the block to be predicted into at least one prediction area according to the size information of the block to be predicted;
[0122] Optionally, the size information of the block to be predicted may be determined, and the block to be predicted may be divided into at least one prediction region according to a division rule corresponding to the size information. The division rule corresponding to the size information may be a pre-set division rule, such as equal-proportion division or division according to a certain size distance interval.
[0123] Optionally, the size of the block to be predicted can be determined based on the block partition signaling, and the block to be predicted can then be divided into at least one prediction region based on a size partitioning rule agreed upon by the encoding and decoding sides. Optionally, the size partitioning rule can be a pre-set partitioning rule, such as equal-proportion partitioning or partitioning based on a certain size distance interval.
[0124] Optionally, the size of the block to be predicted can be determined from the bitstream, and the block to be predicted can be divided into at least one prediction region based on neighboring blocks, and / or co-located blocks, and / or time domain blocks, and / or blocks of other components corresponding to the block to be predicted. For example, the region division rule for the co-located blocks can be used as the region division rule for the current prediction block. Optionally, the region division rule can be a pre-set rule, such as equal-proportional division.
[0125] Optionally, different size information can be set to correspond to different division methods, for example, the division method corresponding to size information greater than a is set to geometric division, the division method corresponding to size information less than b is set to neural network-based division, and the division method corresponding to size information equal to or less than a and greater than or equal to b is set to query table-based division.
[0126] Optionally, the size information of the block to be predicted can be checked to determine whether the size information is too small, such as if the size information is less than a preset size information threshold. If so, the block to be predicted can be directly used as a prediction area. For example, a block can be divided into thirds based on the size information agreed upon by the encoding and decoding sides.
[0127] Optionally, the block to be predicted that does not meet the size information may be skipped from region division. For example, if a block is 4x4, the block to be predicted may be skipped from region division and directly used as a prediction region.
[0128] For example, as shown in Figure 4, the four prediction regions can be divided according to pre-set partition sizes. If the block to be predicted is U, for (a) in Figure 4, the size (W×H) of the block to be predicted U is 128×64, and the block to be predicted U can be divided into three rectangular prediction regions with sizes of 128×20, 128×17, and 128×27, respectively. Different prediction models can be set for these three prediction regions, such as cross-component prediction models, such as CCCM (convolutional cross-component model), GLM (filter-based linear model), and CCLM (cross-component linear model). Each prediction region can then be predicted using the corresponding cross-component prediction model. For example, CCCM can be used for prediction of a 128×20 prediction region. GLM can be used for prediction of a 128×17 prediction region. CCLM can be used for prediction of a 128×27 prediction region.
[0129] Alternatively, for (b) in Figure 4 , the size (W×H) of the block to be predicted U is 64×128. The block to be predicted U can be divided into three rectangular prediction regions with sizes of 28×128, 16×128, and 28×128, respectively. Different prediction models can be set for these three prediction regions, such as a cross-component prediction model, such as CCCM, CCCM, and CCCM. Subsequently, each prediction region can be predicted using the corresponding cross-component prediction model. For example, CCCM is used for prediction of sizes 28×128, 16×128, and 28×128.
[0130] Optionally, for (c) in Figure 4, the size (W×H) of the block to be predicted U is 128×128, and the block to be predicted U can be divided into four rectangular prediction areas with sizes of 90×79, 38×79, 90×49, and 38×49, respectively. Different prediction models can be set for these four prediction areas, such as cross-component prediction models, such as CCCM, CCLM, GLM, and FLM (Filter-based Linear Model). Subsequently, each prediction area can be predicted using the corresponding cross-component prediction model. For example, CCCM can be used for prediction of a prediction area of 90×79. GLM can be used for prediction of a prediction area of 90×49. CCLM can be used for prediction of a prediction area of 38×79. FLM can be used for prediction of a prediction area of 38×49.
[0131] Optionally, for (d) in FIG4 , the size (W×H) of the block to be predicted U is 128×128, and the block to be predicted U can be divided into two non-rectangular prediction areas, and CCLM can be used to perform prediction on these two non-rectangular prediction areas.
[0132] Optionally, by dividing the block to be predicted based on the size information of the block to be predicted to obtain at least one prediction area, it is possible to avoid the phenomenon that the divided prediction area is restricted by a power of 2, and the shape of the divided prediction area is not restricted to a regular shape, that is, it can be an arbitrary shape, so that the block to be predicted can be divided in the prediction stage for the subsequent prediction process. By predicting the prediction area, the prediction accuracy can be improved, thereby improving the efficiency of video encoding and / or decoding.
[0133] Method 2: dividing the block to be predicted into at least one prediction area according to neighboring pixels of the block to be predicted;
[0134] Optionally, the neighbor pixels may be pixels from neighbor blocks, pixels from other components of the block to be predicted, pixels of non-neighbor blocks and / or co-located blocks, or pixels of a time domain block corresponding to the block to be predicted.
[0135] Optionally, the time domain block may be a block divided in the time domain.
[0136] Optionally, if the block to be predicted is a block of the second component, the neighboring pixels of the block to be predicted may include neighboring pixels of the block of the first component and / or neighboring pixels of the block of the second component corresponding to the block to be predicted.
[0137] Alternatively, the neighboring pixels may include pixels of blocks located above, to the left, to the upper left, to the upper right, or to the lower left of the block of the second component and / or the block of the first component.
[0138] Alternatively, in the dual-tree mode, the neighboring pixels may include pixels of blocks located below, to the right, or to the right below the block of the second component and / or the block of the first component.
[0139] Optionally, the second component may be any component of the YUV components, and the third component is different from the second component and is also one of the YUV components.
[0140] Optionally, a lookup table may be set, in which different pixel intervals are set to correspond to different division methods.
[0141] Optionally, the lookup table may be queried to determine the pixel interval to which the neighboring pixels of the block to be predicted belong, determine a division method corresponding to the pixel interval, and divide the block to be predicted according to the division method to obtain at least one prediction area.
[0142] Optionally, the division method may include division depth and / or division times, the shape of the prediction area obtained after division, etc.
[0143] Optionally, a preset rule may be set such that when the neighboring pixels of the block to be predicted meet the preset rule, the block to be predicted is divided into a preset number of prediction regions, such as three prediction regions. Alternatively, the preset rule may be to set a division method based on the number of neighboring pixels, or to divide the block to be predicted based on the pixel values of the neighboring pixels, such as not performing a division when the number of neighboring pixels is less than a certain threshold. Alternatively, the preset rule may be to match the neighboring pixels with a plurality of preset neighboring pixel intervals, determine a matching neighboring pixel interval that matches the neighboring pixel, and determine division information corresponding to the matching neighboring pixel interval (such as a division direction, number of regions, etc.), and then divide the block to be predicted into at least one prediction region based on the division information. Alternatively, the division information corresponding to each neighboring pixel interval may be set in advance, and the division information corresponding to each neighboring pixel interval may be different. Alternatively, when a neighboring pixel belongs to a certain neighboring pixel interval, the neighboring pixel is determined to match the neighboring pixel interval. Alternatively, the preset rule may be a rule set by the user based on scenario requirements.
[0144] Optionally, by dividing the block to be predicted based on its neighboring pixels to obtain at least one prediction region, the neighboring pixels of the neighboring blocks can be comprehensively considered during the division to improve the effectiveness of the prediction region obtained by the division. Furthermore, the block to be predicted can be divided during the prediction phase for subsequent prediction. By predicting the prediction region, prediction accuracy can be improved, thereby improving the efficiency of video encoding and / or decoding.
[0145] Method three, dividing the block to be predicted into at least one prediction area according to the division information of the block to be predicted;
[0146] Optionally, the partitioning information agreed upon by the encoding side and the decoding side may be obtained, and the partitioning information of the block to be predicted may be determined based on the agreed partitioning information, for example, by directly using the agreed partitioning information as the partitioning information of the block to be predicted. Alternatively, the agreed partitioning information may be transformed to obtain the partitioning information of the block to be predicted.
[0147] Optionally, the partition information of the block to be predicted may be pre-set default partition information, partition information determined according to preset rules, or partition information input and set by the user. Optionally, the preset rules may refer to the preset rules in the above-mentioned method 2. The preset rules are not limited to the preset rules in the above-mentioned method 2 and are not limited here.
[0148] Optionally, the division information may include the number of divisions, and / or the division method, and / or the division order, etc.
[0149] Optionally, after determining the partition information of the prediction block, the block to be predicted can be directly partitioned. For example, if the partition information is to divide the block into 2 or 3 parts according to a width:height ratio of 1:2, 1:3, 2:1, or 3:1, then the block to be predicted can be divided into 2 or 3 prediction areas according to a width:height ratio of 1:2, 1:3, 2:1, or 3:1. For example, if the partition information includes dividing the block into two non-rectangular areas using a method similar to geometric partitioning, then the block to be predicted can be divided using a method similar to geometric partitioning, and optionally, the areas obtained by the partitioning include two non-rectangular areas, and at least one area obtained by the partitioning is used as the prediction area.
[0150] Optionally, by dividing the block to be predicted based on the division information of the block to be predicted to obtain at least one prediction area, the block to be predicted can be divided in the prediction stage for subsequent prediction process. By predicting the prediction area, the prediction accuracy can be improved, thereby improving the efficiency of video encoding and / or decoding.
[0151] Method 4: dividing the second component and / or the third component of the block to be predicted into at least one prediction area;
[0152] Optionally, each image block may include three components: YUV components, RGB components, or other components. In this embodiment, only YUV components are used as an example. For example, as shown in FIG5 , there is an image frame composed of images of Y, U, and V components.
[0153] Optionally, as shown in FIG6 , an image including a Y component is included; as shown in FIG7 , an image including a U component is included; and as shown in FIG8 , an image including a V component is included.
[0154] Optionally, the first component may be a Y component, the second component may be a U component, and the third component may be a V component.
[0155] Optionally, the block to be predicted may be divided into at least one prediction area according to the first component, the second component and / or the third component of the block to be predicted.
[0156] Optionally, the second component and / or the third component of the block to be predicted may be divided into at least one prediction region according to the division information and / or division method of the second component of the block to be predicted.
[0157] Optionally, the second component and / or the third component of the block to be predicted may be divided into at least one prediction area according to size information of the block to be predicted.
[0158] Optionally, the second component and / or the third component of the block to be predicted may be divided into at least one prediction region according to size information of neighboring blocks and / or co-located blocks of the block to be predicted.
[0159] Optionally, the second component and / or the third component of the block to be predicted is divided into at least one prediction area according to neighboring pixels of the block to be predicted.
[0160] Optionally, the second component and / or the third component of the block to be predicted may be divided into at least one prediction area according to non-neighboring pixels of the block to be predicted.
[0161] Optionally, the second component and / or the third component of the block to be predicted is divided into at least one prediction area according to division information of the block to be predicted.
[0162] Optionally, the second component and / or the third component of the block to be predicted may be divided into at least one prediction region according to the division information of the first component in the block to be predicted.
[0163] Optionally, the third component of the block to be predicted may be divided into at least one prediction region according to the division information of the first component and the division information of the second component of the block to be predicted.
[0164] Optionally, the second component of the block to be predicted may be divided into at least one prediction region according to the division information of the first component and the division information of the third component of the block to be predicted.
[0165] Optionally, the second component and / or the third component of the block to be predicted is divided into at least one prediction area according to pixel gradient values of neighboring blocks.
[0166] Optionally, the second component and / or the third component of the block to be predicted may be divided into at least one prediction area according to pixel gradient values of non-neighboring blocks and / or co-located blocks.
[0167] Optionally, the prediction region may include a second component prediction region and / or a third component prediction region.
[0168] Optionally, by dividing the second component and the third component of the block to be predicted into at least one prediction area when dividing the block to be predicted, it is possible to divide the different components of the block to be predicted, and then to predict the prediction areas of different components, thereby improving the prediction accuracy and improving the efficiency of video encoding and / or decoding.
[0169] Method 5: Divide the block to be predicted into at least one prediction area according to pixel gradient values of neighboring blocks.
[0170] Optionally, pixel gradient values may be calculated based on pixels of neighboring blocks, and the block to be predicted may be divided into regions based on the pixel gradient values. For example, a dividing line may be determined based on the magnitude of the pixel gradient values in the horizontal and / or vertical directions, and the block to be predicted may be divided based on the dividing line.
[0171] Optionally, pixel gradient values may be calculated based on pixels of non-neighboring blocks and pixels of neighboring blocks, or based on pixels of co-located blocks, and then the block to be predicted may be divided based on the pixel gradient values to obtain at least one prediction area.
[0172] For example, as shown in FIG9 , the prediction area is divided according to the pixel gradient value calculation. The division can be performed according to a pre-set division size. The corresponding division information can be obtained by calculating the neighboring pixels of the block to be predicted, and the division is performed based on the division information.
[0173] Optionally, if the block to be predicted is the current block, reconstructed neighboring pixels of the Cb component of the current block U may be obtained, such as the pixels shown in the dotted L-shaped area in the image corresponding to a in FIG9 .
[0174] Optionally, the neighboring pixels may include above, left, upper-left, lower-left, and upper-right pixels of the current block.
[0175] Optionally, a in FIG9 is the current block U cb and its neighboring pixels, and then after gradient calculation, the current block U as shown in b in Figure 9 is obtained cb and its neighbor pixel gradients, and the region is segmented according to the neighbor pixel gradients to obtain the predicted region shown in Figure 9 c, such as U cb1 、U cb2 、U cb3 and U cb4 .
[0176] Optionally, the Sobel operator may be used to calculate the gradient of the L-shaped region, that is, the pixel gradient value of the neighboring block, to obtain an image corresponding to b in FIG9 .
[0177] Optionally, darker colors in an image represent larger gradients and faster pixel changes.
[0178] Optionally, the x-direction operator G of the Sobel operator x The definition is as follows:
[0179] Optionally, the y-direction operator G of the Sobel operator y The definition is as follows:
[0180] Alternatively, the total gradient calculation can be or
[0181] Optionally, I represents the pixels of the original image, G x and G y represent the horizontal and vertical gradient operators respectively.
[0182] Optionally, the Scharr operator may be used for calculation.
[0183] Optionally, multiple prediction regions may be obtained according to a preset division rule. For example, a target number of prediction regions corresponding to the block to be predicted may be determined according to the preset division rule, and the block to be predicted may be divided into the target number of prediction regions. Optionally, the division rule may be the same as the aforementioned division rule, such as dividing the block to be predicted according to a certain size ratio to obtain multiple prediction regions.
[0184] Optionally, by dividing the block to be predicted based on the gradient values of neighboring pixels of the block to be predicted to obtain at least one prediction region, the gradient values of neighboring pixels of neighboring blocks can be comprehensively considered during the division process to improve the effectiveness of the prediction region obtained by the division. Furthermore, the block to be predicted can be divided during the prediction phase for subsequent prediction. By predicting the prediction region, prediction accuracy can be improved, thereby improving the efficiency of video encoding and / or decoding.
[0185] Optionally, before performing prediction based on the prediction model and / or model parameters corresponding to the prediction area, the prediction model and / or model parameters corresponding to each prediction area may be determined.
[0186] Optionally, the prediction model includes a cross-component prediction model. The prediction model may be a prediction model corresponding to the prediction region itself, or a prediction model corresponding to another region. Optionally, the other region may be a region at the same level as the prediction region, such as a region after the prediction block is divided. Alternatively, the other region may be a region at a different level than the prediction region, such as a prediction model corresponding to a neighboring block or a prediction model corresponding to a non-neighboring block.
[0187] Optionally, different prediction regions may use the same or different cross-component prediction models. Optionally, the cross-component prediction model used by the prediction region is obtained based on the reference pixels of the prediction region. Optionally, model parameters of the cross-component prediction model are calculated based on the reference pixels of the prediction region.
[0188] Optionally, the cross-component prediction model and / or model parameters used in the prediction region are obtained in the bitstream. Optionally, the cross-component prediction model and / or model parameters used in the second component prediction region can be obtained based on the reference pixels of the second component prediction region and / or the reference pixels of the third component prediction region.
[0189] Optionally, the reference pixels may be composed of target neighboring pixels of the first component of the current block, and / or the second component of the current block, and / or the third component of the current block.
[0190] Optionally, the target neighbor pixels may include pixels above, to the left, to the upper left, to the upper right, or to the lower left of the block of the same component.
[0191] Alternatively, in dual-tree mode, the target neighbor pixels may include pixels below, to the right, and to the lower right of a block of the same component. For example, in dual-tree mode, if there are luma neighbor pixels to the right and below of a luma component block, the block has been reconstructed. The luma neighbor pixels to the right and below of the luma component block may be selected to determine the reference pixel.
[0192] Optionally, for each prediction area, prediction may be performed using a prediction model and / or model parameters corresponding to the prediction area.
[0193] Optionally, for each prediction region, a prediction mode corresponding to the prediction region may be determined, and a prediction model and / or model parameters to be applied may be determined based on the prediction mode.
[0194] Optionally, the prediction mode may include an inter-component prediction mode, a temporal prediction mode, an intra-frame and / or an inter-frame prediction mode.
[0195] Optionally, the prediction mode may further include a mode for performing prediction based on at least one of CCLM, FLM, GLM, CCCM, GL-CCCM, WCP (Worst-Case Perturbations, adversarial model), etc.
[0196] Optionally, the time domain prediction mode may be a prediction mode in the time domain, that is, a prediction mode different from the current frame / time, such as an adjacent block prediction mode and / or a non-adjacent block prediction mode of the co-located block of the current block in a reference frame or a collocated picture.
[0197] Optionally, the temporal prediction mode may be derived from the prediction mode of a neighboring frame of the frame containing the current block. The neighboring frame may be an intra-frame coded frame or a mixed intra-frame or inter-frame coded frame. For example, if the frame containing the current block is a P-frame (unidirectionally predicted frame) or a B-frame (bidirectionally predicted frame), and the GOP (Group of Picture) number is n, the neighboring frame GOP number may be n-1, n-2, ..., n+1, n+2, etc., and may be an I-frame (keyframe, intra-frame coded frame), a P-frame, or a B-frame.
[0198] Optionally, the temporal prediction mode is an intra-frame prediction mode, such as a cross-component intra-frame prediction mode for a co-located block in a collocated frame. For the current block, there may be multiple temporal prediction modes, and they may be sorted in a certain order. For example, when a co-located block in a collocated frame uses an inter-frame prediction mode, the cross-component intra-frame prediction mode of a neighboring block of the co-located block is used instead.
[0199] Optionally, if the current block is offset from the co-located block in the collocated frame, the offset target block (i.e., the offset co-located block) can be obtained based on the corresponding motion vector, and the cross-component intra-frame prediction mode of the target block can be used as the time domain prediction mode of the current block.
[0200] Optionally, this embodiment divides the block to be predicted into at least one prediction area; and performs prediction based on the prediction model and / or model parameters corresponding to at least one prediction area. Through the above technical solution, it is possible to divide the block to be predicted into multiple prediction areas during the prediction stage, thereby avoiding the phenomenon that the size of each prediction area is limited by a power of 2, and / or the prediction area is rectangular, and the prediction is performed based on the prediction model and / or model parameters corresponding to the prediction area, such as performing cross-component prediction based on a cross-component prediction model, thereby avoiding the phenomenon that all pixels of the block are predicted using a fixed cross-component prediction model and model parameters, resulting in low prediction accuracy, thereby improving the prediction accuracy of the block prediction, and thereby improving the efficiency of video encoding and / or decoding.
[0201] Second embodiment
[0202] Based on the above-mentioned first embodiment, a second embodiment is proposed.
[0203] In this embodiment, referring to FIG. 10 , step S20 includes the following steps:
[0204] S21, performing cross-component prediction on at least one prediction area according to the prediction model and / or model parameters corresponding to the at least one prediction area, and determining or generating a prediction result;
[0205] Optionally, for each prediction region, a prediction model and / or model parameters corresponding to the prediction region are first determined.
[0206] Optionally, the prediction models and / or model parameters corresponding to each prediction area may be the same or different.
[0207] Optionally, if the prediction type of at least one prediction region is cross-component prediction, the at least one prediction region may be predicted based on the cross-component prediction model and / or model parameters corresponding to the at least one prediction region to determine or generate a prediction result.
[0208] Optionally, each prediction region corresponds to a prediction result, and the prediction method for each prediction region can be the same or different, for example, one prediction region performs cross-component prediction, and another prediction region performs same-component prediction.
[0209] S22: Determine or generate a prediction block according to the prediction results of the plurality of prediction areas.
[0210] Optionally, the prediction results of the obtained multiple prediction areas are weightedly summed and / or combined and / or spliced to determine or generate a prediction block.
[0211] Optionally, the weighted summation and / or combination or splicing may include prediction results of multiple prediction areas, such as combining pixels of the prediction areas according to the order of the prediction areas in the image, to determine or generate a prediction block.
[0212] In this embodiment, cross-component prediction is performed on at least one prediction area based on the prediction model and / or model parameters corresponding to the at least one prediction area to determine or generate a prediction result, and weighted summation and / or combination or splicing are performed based on multiple prediction results to determine or generate a prediction block. It is then possible to predict each prediction area and then determine or obtain a prediction block based on the prediction result, thereby improving the prediction accuracy and thereby improving the efficiency of video encoding and / or decoding.
[0213] Third embodiment
[0214] Based on any of the above embodiments, a third embodiment is proposed.
[0215] In this embodiment, step S20 may be at least one of the following:
[0216] Method six, determining or obtaining a second component prediction block according to prediction results of multiple second component prediction areas;
[0217] Optionally, if the second component of the prediction block is divided to obtain multiple second component prediction areas, the second component prediction block is determined or obtained based on the result of weighted summation and / or combination and / or splicing of the prediction results of the multiple second component prediction areas.
[0218] Optionally, the first component prediction area corresponding to each second component prediction area can be determined, and the prediction results of multiple second component prediction areas can be weighted summed up and / or combined and / or spliced according to the results of weighted summation and / or combination and / or splicing of each first component prediction area to determine or obtain the second component prediction block.
[0219] Optionally, when performing weighted summation and / or combination and / or splicing of the prediction results, the prediction results of the second component prediction area can be weighted summation and / or combination and / or splicing to determine or obtain the second component prediction block, and then the prediction results of the same component can be weighted summation and / or combination and / or splicing to determine or obtain the second component prediction block, so that the second component prediction block can be directly reconstructed subsequently, thereby improving the efficiency of video encoding and / or decoding.
[0220] Method seven: determine or obtain a third component prediction block according to prediction results of multiple third component prediction areas.
[0221] Optionally, if the third component of the prediction block is divided to obtain multiple third component prediction areas, the third component prediction block is determined or obtained based on the result of weighted summation and / or combination and / or splicing of the prediction results of the multiple third component prediction areas.
[0222] Optionally, the first component prediction area corresponding to each third component prediction area can be determined, and the prediction results of multiple third component prediction areas can be weighted summed up and / or combined and / or spliced according to the results of weighted summation and / or combination and / or splicing of each first component prediction area to determine or obtain the third component prediction block.
[0223] Optionally, when performing weighted summation and / or combination and / or splicing of the prediction results, the prediction results of the third component prediction area can be weighted summation and / or combination and / or splicing to determine or obtain a third component prediction block, and then the prediction results of the same component can be weighted summation and / or combination and / or splicing to obtain the third component prediction block, so that the third component prediction block can be directly reconstructed subsequently, thereby improving the efficiency of video encoding and / or decoding.
[0224] Fourth embodiment
[0225] Based on any of the above embodiments, a fourth embodiment is proposed.
[0226] In this embodiment, the processing method further includes at least one of the following methods 8 to 11:
[0227] Mode eight, reconstructing the block based on the prediction residual obtained from the bitstream and the second component prediction block, and determining or generating a reconstruction of the block;
[0228] Optionally, at the decoding side, if the prediction results of multiple second component prediction areas of the prediction block are weighted summed and / or combined and / or concatenated to obtain the second component prediction block, the corresponding information can be extracted from the bitstream and decoded to obtain the prediction residual of the block.
[0229] Alternatively, corresponding information may be extracted from the bitstream and decoded to obtain a parameter or syntax element, which is then transformed according to a transformation rule agreed upon with the encoding side to obtain a prediction residual for the block. Optionally, the transformation rule may be addition, subtraction, or format conversion of the parameter or syntax element.
[0230] Optionally, the prediction residual and the second component prediction block can be input into a preset adder, which calculates the prediction residual and the second component prediction block, and clips the adder result so that the clipped result is within a legal pixel value range to complete the reconstruction operation and generate a reconstruction of the block. Optionally, the legal pixel value range can be a legal pixel range of 0-255.
[0231] Optionally, the prediction residual and the second component prediction block may be added together by an adder, and the result of the adder may be clipped so that the clipped result is within a legal pixel value interval to complete the reconstruction operation.
[0232] Optionally, the reconstruction of the block can be determined or generated by reconstructing the prediction residual and the second component prediction block obtained based on the code stream, thereby ensuring that the reconstruction of the generated block is consistent with the block of the second component in the encoding side, thereby improving the efficiency of video encoding and / or decoding.
[0233] Mode nine, reconstructing the block based on the prediction residual obtained from the bitstream and the third component prediction block, and determining or generating a reconstruction of the block;
[0234] Optionally, at the decoding side, if the prediction results of multiple third component prediction areas of the prediction block are weighted summed and / or combined and / or concatenated to obtain the third component prediction block, the corresponding information can be extracted from the bitstream and decoded to obtain the prediction residual.
[0235] Alternatively, corresponding information may be extracted from the bitstream and decoded to obtain a parameter or syntax element, which is then transformed according to a transformation rule agreed upon with the encoding side to obtain a prediction residual for the block. Optionally, the transformation rule may be addition, subtraction, or format conversion of the parameter or syntax element.
[0236] Optionally, the prediction residual and the third component prediction block can be input into a preset adder, the prediction residual and the third component prediction block are calculated by the adder, and the result of the adder is clipped so that the clipped result is within a legal pixel value range to complete the reconstruction operation and generate a reconstruction of the block. Optionally, the legal pixel value range can be a legal pixel range of 0-255.
[0237] Optionally, the prediction residual and the third component prediction block may be added together by an adder, and the result of the adder may be clipped so that the clipped result is within a legal pixel value interval to complete the reconstruction operation.
[0238] For example, as shown in FIG11 , the Cb component is predicted by the Y component. Optionally, the parsing module performs entropy decoding, inverse quantization and inverse transformation on the bit stream b to obtain the prediction residual res (U cb ). Optionally, after parsing, reconstruction and bit stream b, the Y component is obtained, that is, rec(U y ). Optionally, the Cr component is predicted by the Y component, or the Cr component is predicted by the Cb component.
[0239] Optionally, according to a preset division rule, the component to be predicted U cb Divide into n prediction areas, where n is a positive integer greater than 1. For example, as shown in Figure 11, U cb,1 ,...,U cb,n Optionally, the division rule may include at least one of approximately uniform horizontal division, approximately uniform vertical division, oblique line division, unequal division, division calculated based on the gradient of neighboring pixels of the block, etc. For any prediction area U cb,i , the reference pixel set S in the Y component and / or Cb component can be obtained according to the preset rules cb,i For example, S cb,1 ,...,S cb,n .
[0240] Optionally, the preset rule may be to classify the reference pixels of the Y component into corresponding preset areas of the U component according to the principle of closest distance.
[0241] Optionally, when determining the prediction area U cb,i and the reference pixel set S cb,i In the case of , the model parameters of the cross-component prediction model can be calculated by the prediction module as shown in Figure 11. Then, the calculated model parameters are used to reconstruct the luma classification rec(U y ) is input, and the prediction is pred(U cb,i ), and then get pred(U cb,1 ),...,pred(U cb,n ). Then input it into the merging module to get the complete pred(U cb ).
[0242] Optionally, pred(U cb ) is input to the adder, and the adder is pred(U cb ) and the prediction residual res(Ucb ) build and add, get U cb Reconstruction, that is, rec(U cb ). The prediction residual is transformed, quantized and entropy coded to obtain the bit stream b.
[0243] Optionally, U can be the block currently being encoded and / or decoded, such as the CU in the H.266 / VCC protocol. y , U cb , U cr , which can be the three components of block U, such as brightness / luma, blue chroma / cb, and red chroma / cr.
[0244] Alternatively, org(U) can be the original pixels of block U. Similarly, org(U x ) can be the original pixel of the x component of block U. pred(U) can be the prediction of block U. Similarly, pred(U x ) can be the prediction of the x component of block U. rec(U) can be the reconstruction of block U. Similarly, rec(U x ) can be the reconstruction of the x component of block U. res(U) can be the residual of block U. Similarly, res(U x ) can be the residual of the x component of block U. x,i Can be block U x,i The x-component of the i-th prediction region. x It can be the x component U of the current block U x The reference pixel set is the reconstructed pixel. b can be a video bitstream / codestream.
[0245] Optionally, the reconstruction of the block can be determined or generated by reconstructing the prediction residual and the third component prediction block obtained based on the code stream, thereby ensuring that the reconstruction of the generated block is consistent with the block of the third component in the encoding side, thereby improving the efficiency of video encoding and / or decoding.
[0246] Method 10: determining or generating a prediction residual according to the pixels of the block to be predicted and the second component prediction block;
[0247] Optionally, on the encoding side, if the prediction results of multiple second component prediction areas of the to-be-predicted block are weighted summed and / or combined and / or spliced to obtain the second component prediction block, the pixels of the to-be-predicted block can be obtained, and then the pixels of the to-be-predicted block and the second component prediction block can be input into a preset subtractor to perform calculations on the pixels of the to-be-predicted block and the second component prediction block through the subtractor, such as subtracting the pixels of the to-be-predicted block from the pixels of the second component prediction block to obtain a residual, which can be used as a prediction residual, and the prediction residual is transformed, quantized, and entropy encoded to obtain a bitstream b, that is, the prediction residual is encoded into the bitstream.
[0248] Optionally, a transformation rule agreed upon with the decoding side can be determined, and the residual can be transformed according to the transformation rule to obtain a prediction residual, which can then be encoded into the bitstream. Optionally, the pixels of the block to be predicted can be the original pixels of the block to be predicted. Optionally, the transformation rule can include a rule for format conversion of the residual, and can also be configured and adjusted according to user needs.
[0249] Optionally, by determining or generating a prediction residual based on the pixels of the block to be predicted and the second component prediction block, the phenomenon of excessive signaling consumption caused by directly encoding the second component prediction block is avoided, thereby improving the efficiency of video encoding and / or decoding.
[0250] In an eleventh manner, a prediction residual is determined or generated according to pixels of the block to be predicted and the third component prediction block.
[0251] Optionally, on the encoding side, if the prediction results of multiple third component prediction areas of the to-be-predicted block are weighted summed and / or combined and / or spliced to obtain the third component prediction block, the pixels of the to-be-predicted block can be obtained, and then the pixels of the to-be-predicted block and the third component prediction block can be input into a preset subtractor to perform calculations on the pixels of the to-be-predicted block and the third component prediction block through the subtractor, such as subtracting the pixels of the to-be-predicted block from the pixels of the third component prediction block to obtain a residual, which can be used as a prediction residual, and the prediction residual is transformed, quantized, and entropy encoded to obtain a bitstream b, that is, the prediction residual is encoded into the bitstream.
[0252] Alternatively, a transformation rule agreed upon with the decoding side can be determined, and the residual is transformed according to the transformation rule to obtain a prediction residual, which is then encoded into the bitstream. Optionally, the transformation rule can include a rule for formatting the residual, and can also be set and adjusted according to user needs.
[0253] Optionally, the pixels of the block to be predicted may be original pixels of the block to be predicted.
[0254] For example, as shown in FIG12 , the Cb component is predicted by the Y component of the block to be predicted. Optionally, the Y component has completed the encoding process, that is, it is known that rec(U y ). Optionally, the Cr component is predicted by the Y component, or the Cr component is predicted by the Cb component.
[0255] Optionally, according to a preset division rule, the component to be predicted U cb Divide into n prediction areas, where n is a positive integer greater than 1. For example, as shown in Figure 12, U cb,1 ,...,U cb,nOptionally, the division rule may include at least one of approximately uniform horizontal division, approximately uniform vertical division, oblique line division, unequal division, division calculated based on the gradient of neighboring pixels of the block, etc. For any prediction area U cb,i , the reference pixel set S in the Y component and / or Cb component can be obtained according to the preset rules cb,i For example, S cb,1 ,...,S cb,n .
[0256] Optionally, the preset rule may be to classify the reference pixels of the Y component into corresponding preset areas of the U component according to the principle of closest distance.
[0257] Optionally, when determining the prediction area U cb,i and the reference pixel set S cb,i In the case of , the model parameters of the cross-component prediction model can be calculated by the prediction module as shown in Figure 12. Then, the calculated model parameters are used to reconstruct the luma classification rec(U y ) is input, and the prediction is pred(U cb,i ), and then get pred(U cb,1 ),...,pred(U cb,n ). Then input it into the merging module to get the complete pred(U cb ). Optionally, pred(U cb ) is input to the subtractor, and the subtractor is used to generate org(U cb ) minus pred(U cb ), and get the prediction residual res(U cb ). The prediction residual is transformed, quantized and entropy coded to obtain a bit stream b. Optionally, the specific parameters can refer to the description of Figure 11 in the above method 10.
[0258] Optionally, by determining or generating a prediction residual based on the pixels of the block to be predicted and the third component prediction block, the phenomenon of excessive signaling consumption caused by directly encoding the third component prediction block is avoided, thereby improving the efficiency of video encoding and / or decoding.
[0259] Fifth embodiment
[0260] Based on any of the above embodiments, a fifth embodiment is proposed.
[0261] In this embodiment, the prediction region includes a second component prediction region and / or a third component prediction region.
[0262] Optionally, the processing method further includes at least one of the following:
[0263] The second component prediction area includes a rectangular area and / or a non-rectangular area;
[0264] The third component prediction area includes a rectangular area and / or a non-rectangular area.
[0265] Alternatively, the second component may be any one of a luminance component, a blue chrominance component, and a red chrominance component, and the third component may be any one of the luminance component, the blue chrominance component, and the red chrominance component that is different from the second component.
[0266] Optionally, when dividing the block to be predicted into at least one prediction region, if the division is performed on the second component of the block to be predicted, the obtained prediction region includes the second component prediction region. If the division is performed on the third component of the block to be predicted, the obtained prediction region includes the third component prediction region.
[0267] Optionally, when dividing the second component and / or the third component of the block to be predicted, the division rule is not limited, and may be approximately uniform horizontal division, approximately uniform vertical division, diagonal division, or uneven division. Therefore, the shape and size of the second component prediction region obtained by division are not fixed, that is, it may be a rectangular region or a non-rectangular region.
[0268] Method 12: determining or obtaining a prediction model and / or model parameters corresponding to at least one prediction area based on reference pixels of the block to be predicted;
[0269] Optionally, after determining the reference pixels of the block to be predicted, the reference pixels may be input into a preset model module, and a prediction model and / or model parameters corresponding to at least one prediction area may be obtained as output.
[0270] Optionally, if there are multiple prediction regions and multiple reference pixels, for each prediction region, a block closest to the prediction region among the blocks corresponding to each reference pixel may be determined, and the prediction model and / or model parameters used by the closest block may be selected as the prediction model and / or model parameters corresponding to the prediction region.
[0271] Alternatively, a lookup table may be provided that indicates prediction models and / or model parameters corresponding to reference pixel intervals based on preset rules, and the lookup table may be queried based on the reference pixels of the block to be predicted to determine or obtain the prediction model and / or model parameters corresponding to at least one prediction region. Alternatively, the preset rules may refer to any of the aforementioned embodiments, and the specific configuration of the preset rules is not limited herein.
[0272] Optionally, the prediction model and / or model parameters corresponding to at least one prediction area can be determined or obtained based on the reference pixels of the block to be predicted, thereby avoiding the phenomenon that only fixed prediction models and / or model parameters can be used when predicting the block to be predicted, improving the prediction accuracy of block prediction, and thus improving the efficiency of video encoding and / or decoding.
[0273] Method 13: determining or obtaining a cross-component prediction model and / or model parameters of the second component prediction region based on reference pixels of the second component prediction region and / or reference pixels of the third component prediction region;
[0274] Optionally, after determining the reference pixels of the second component prediction area and / or the reference pixels of the third component prediction area of the block to be predicted, the reference pixels of the second component prediction area and / or the reference pixels of the third component prediction area can be input into a preset model module, and the cross-component prediction model and / or model parameters of the second component prediction area are output.
[0275] Optionally, a lookup table may be provided that indicates prediction models and / or model parameters corresponding to reference pixel intervals according to preset rules, and the lookup table may be queried based on the reference pixels of the second component prediction region and / or the reference pixels of the third component prediction region to determine or obtain the cross-component prediction model and / or model parameters for the second component prediction region. Optionally, the preset rules may refer to any of the above embodiments, and the specific configuration of the preset rules is not limited herein.
[0276] Optionally, the cross-component prediction model and / or model parameters of the second component prediction region may be determined or obtained based on the reference pixels of the first component and the reference pixels of the second component prediction region and / or the third component prediction region.
[0277] Optionally, the cross-component prediction model and / or model parameters of the second component prediction area can be determined or obtained based on the reference pixels of the second component prediction area and / or the reference pixels of the third component prediction area, thereby avoiding the phenomenon that only fixed prediction models and / or model parameters can be used when predicting the prediction block, improving the prediction accuracy of block prediction, and thus improving the efficiency of video encoding and / or decoding.
[0278] In a fourteenth method, a cross-component prediction model and / or model parameters of the third component prediction region are determined or obtained based on the reference pixels of the second component prediction region and / or the reference pixels of the third component prediction region.
[0279] Optionally, after determining the reference pixels of the second component prediction area and / or the reference pixels of the third component prediction area of the block to be predicted, the reference pixels of the second component prediction area and / or the reference pixels of the third component prediction area can be input into a preset model module, and the cross-component prediction model and / or model parameters of the third component prediction area are output.
[0280] Optionally, a lookup table may be provided that corresponds to a prediction model and / or model parameters for reference pixel intervals according to a preset rule, and the lookup table may be queried based on the reference pixels of the second component prediction region and / or the reference pixels of the third component prediction region to determine or obtain the cross-component prediction model and / or model parameters for the third component prediction region. Optionally, the preset rules may refer to any of the above embodiments, and the specific configuration of the preset rules is not limited herein.
[0281] Optionally, the cross-component prediction model and / or model parameters of the third component prediction region may be determined or obtained based on the reference pixels of the first component and the reference pixels of the second component prediction region and / or the third component prediction region.
[0282] For example, the prediction model is CCCM. When the prediction area is predicted based on CCCM, as shown in FIG13 , the prediction area U cb , 1 performs CCCM prediction. Optionally, obtain U cb,1 The reference pixels include TL(y), T1(y), L1(y), TL(cb), T1(cb), and L1(cb). Optionally, TL(y)->TL(cb); T1(y)->T1(cb); and L1(y)->L1(cb) can be considered as three sets of mappings to calculate the model parameters of CCCM. The prediction formula of the CCCM model is as follows: predChromaVal = c0*C+c1*N+c2*S+c3*E+c4*W+c5*P+c6*B;
[0283] Optionally, c0-c6 are 7-tap filter parameters. C represents the luminance sample at the corresponding position of the current chrominance sample, and N, S, E, and W are the adjacent samples of the current luminance sample respectively. P is a nonlinear term. Optionally, P = (C*C+midVal)>>bitDepth; B = midVal, and is a bias term, representing a scalar offset between input and output, and is set to the intermediate chrominance value. Optionally, for 10-bit video, B = 512. Optionally, according to the prediction formula of the model CCCM, minimize the error between the y to cb prediction and the reconstructed Cb, that is, it can be performed according to the following formula: MSE = E[(pred(Cb)-rec(Cb)) 2 ];
[0284] Alternatively, taking TL(y)->TL(cb) as an example, we can get The values of model parameters c0-c6 can be obtained by LDL decomposition. y,1 And model parameters, such as N, S, E, W, and C are used to predict the Cb component and obtain the U in the Cb component. cb,1 predictions.
[0285] Optionally, the cross-component prediction model and / or model parameters of the third component prediction area can be determined or obtained based on the reference pixels of the second component prediction area and / or the reference pixels of the third component prediction area, thereby avoiding the phenomenon that only fixed prediction models and / or model parameters can be used when predicting the prediction block, improving the prediction accuracy of block prediction, and thus improving the efficiency of video encoding and / or decoding.
[0286] In an optional embodiment, the processing method further includes Mode 15 and / or Mode 16:
[0287] In a fifteenth embodiment, the reference pixels of the second component prediction area include neighboring pixels of the first component of the block to be predicted and neighboring pixels of the second component of the block to be predicted;
[0288] Optionally, the neighboring pixels of the first component of the block to be predicted may include pixels of neighboring blocks above, to the left, to the upper left, to the upper right, and to the lower left of the first component of the block to be predicted. Furthermore, when in dual-tree mode, the neighboring pixels of the first component of the block to be predicted may also include pixels of neighboring blocks below, to the right, and to the lower right of the first component of the block to be predicted.
[0289] Optionally, the neighboring pixels of the first component of the block to be predicted may include adjacent block pixels adjacent to the first component of the block to be predicted and non-adjacent non-adjacent block pixels.
[0290] Optionally, the neighboring pixels of the second component of the block to be predicted may include pixels of neighboring blocks above, to the left, to the upper left, to the upper right, and to the lower left of the second component of the block to be predicted. Furthermore, when in dual-tree mode, the neighboring pixels of the neighboring blocks below, to the right, and to the lower right of the second component of the block to be predicted may also be included.
[0291] Optionally, the neighboring pixels of the second component of the block to be predicted may include adjacent block pixels adjacent to the second component of the block to be predicted and non-adjacent non-adjacent block pixels.
[0292] Optionally, the reference pixels of the first component prediction area and / or the reference pixels of the second component prediction area may further include pixels of a co-located block, pixels of a time domain block, etc.
[0293] Optionally, by setting the reference pixels of the second component prediction area to include neighboring pixels of the first component of the block to be predicted and neighboring pixels of the second component of the block to be predicted, the cross-component prediction model and / or model parameters determined based on the reference pixels are made more accurate, thereby improving the prediction accuracy of block prediction and thereby improving the efficiency of video encoding and / or decoding.
[0294] In a sixteenth embodiment, the reference pixels of the third component prediction area include neighboring pixels of the first component of the block to be predicted and neighboring pixels of the third component of the block to be predicted.
[0295] Optionally, the neighboring pixels of the third component of the block to be predicted may include pixels of neighboring blocks above, to the left, to the upper left, to the upper right, and to the lower left of the third component of the block to be predicted. Furthermore, when in dual-tree mode, the neighboring pixels of the third component of the block to be predicted may also include pixels of neighboring blocks below, to the right, and to the lower right of the third component of the block to be predicted.
[0296] Optionally, the neighboring pixels of the third component of the block to be predicted may include adjacent block pixels adjacent to the third component of the block to be predicted and non-adjacent non-adjacent block pixels.
[0297] Optionally, the reference pixels of the third component prediction area may also include pixels of a co-located block, pixels of a time domain block, etc.
[0298] For example, as shown in FIG14 , the predicted component U of the current block U is included cb Optionally, the Cb component to be predicted of the current block U and its four prediction areas U cb,1 ,...,U cb,4 If the reference pixel is determined according to the nearest distance rule, the reference pixel can be as follows:
[0299] Block U cb The Y component reference pixels include: TL(y), T1(y), T2(y), TR(y), L1(y), L2(y), BL(y), and U y .
[0300] Block U cb,1 The Y component reference pixels include: TL(y), T1(y), L1(y), and U y,1 .
[0301] Block U cb,2 The Y component reference pixels include: T2(y), TR(y), and U y,2 .
[0302] Block U cb,3 The Y component reference pixels include: T2(y), BL(y), and U y,3 .
[0303] Block Ucb,4 The Y component reference pixels include: BL(y), TR(y), and U y,4 .
[0304] For the reference pixels of Cb classification, it can be shown as follows:
[0305] Block U cb The Cb component reference pixels include: TL(Cb), T1(Cb), T2(Cb), TR(Cb), L1(Cb), L2(Cb), BL(Cb).
[0306] Block U cb,1 The Cb component reference pixels include: TL(Cb), T1(Cb), L1(Cb).
[0307] Block U cb,2 The Cb component reference pixels include: T2(Cb) and TR(Cb).
[0308] Block U cb,3 The Cb component reference pixels include: T2(Cb), BL(Cb).
[0309] Block U cb,4 The Cb component reference pixels include: BL(Cb) and TR(Cb).
[0310] Optionally, in the block U cb When making predictions, all reference pixels have been reconstructed, and then the block U is reconstructed based on the reference pixels. cb Make predictions.
[0311] For example, as shown in FIG15 , the Cr component to be predicted of the current block U and its four prediction regions U cr,1 ,...,U cr,4 If the reference pixel is determined according to the nearest distance rule, the reference pixel can be as follows:
[0312] Block U cr The Y component reference pixels include: TL(y), T1(y), T2(y), TR(y), L1(y), L2(y), BL(y), B1(y), B2(y), R1(y), R2(y), BR(y), and U y .
[0313] Block U cr,1 The Y component reference pixels include: TL(y), T1(y), L1(y), and U y,1 .
[0314] Block U cr,2 The Y component reference pixels include: T2(y), TR(y), R1(y), and U y,2.
[0315] Block U cr,3 The Y component reference pixels include: L2(y), BL(y), B1(y), and U y,3 .
[0316] Block U cr,4 The Y component reference pixels include: B2(y), BR(y), R2(y), and U y,4 .
[0317] Alternatively, for Cr-classified reference pixels, the following may be used:
[0318] Block U cr The Cr component reference pixels include: TL(cr), T1(cr), T2(cr), TR(cr), L1(cr), L2(cr), BL(cr).
[0319] Block U cr,1 The Cr component reference pixels include: TL(cr), T1(cr), L1(cr).
[0320] Block U cr,2 The Cr component reference pixels include: T2(cr), TR(cr).
[0321] Block U cr,3 The Cr component reference pixels include: L2(cr), BL(cr).
[0322] Block U cr,4 The Cr component reference pixels include: BL(cr), TR(cr).
[0323] Optionally, in the block U cr When making predictions, all reference pixels have been reconstructed, and then the block U is reconstructed based on the reference pixels. cr Make predictions.
[0324] Optionally, in the dual-tree mode, the chrominance component and the luminance component may adopt two different independent tree coding structures.
[0325] Optionally, if block U is at the right edge and bottom edge of a non-CTU, cr When making predictions, U y The pixels of the neighboring blocks to the right, below, and to the lower right have been reconstructed.
[0326] Optionally, by setting the reference pixels of the third component prediction area to include neighboring pixels of the first component of the block to be predicted and neighboring pixels of the third component of the block to be predicted, the cross-component prediction model and / or model parameters determined based on the reference pixels are made more accurate, thereby improving the prediction accuracy of block prediction and thereby improving the efficiency of video encoding and / or decoding.
[0327] The present application also provides a processing device, referring to FIG16 . FIG16 is a schematic diagram of the functional modules of the processing device of the present application, which can be provided in a processing device or be the processing device itself. The processing device of the present application includes:
[0328] A division module A10, configured to divide the block to be predicted into at least one prediction region;
[0329] The prediction module A20 is used to perform prediction based on the prediction model and / or model parameters corresponding to the prediction area.
[0330] Optionally, the prediction model comprises a cross-component prediction model.
[0331] Optionally, dividing the block to be predicted into at least one prediction area includes at least one of the following:
[0332] Dividing the block to be predicted into at least one prediction area according to size information of the block to be predicted;
[0333] Dividing the block to be predicted into at least one prediction area according to neighboring pixels of the block to be predicted;
[0334] Dividing the block to be predicted into at least one prediction area according to the division information of the block to be predicted;
[0335] Dividing the second component and / or the third component of the block to be predicted into at least one prediction area;
[0336] The block to be predicted is divided into at least one prediction area according to pixel gradient values of neighboring blocks.
[0337] Optionally, the prediction module A20 is configured to:
[0338] Performing cross-component prediction on at least one prediction area based on a prediction model and / or model parameters corresponding to at least one prediction area, and determining or generating a prediction result;
[0339] A prediction block is determined or generated according to the prediction results of the plurality of prediction areas.
[0340] Optionally, determining or generating a prediction block according to the prediction results of the plurality of prediction areas includes at least one of the following:
[0341] Determine or obtain a second component prediction block according to prediction results of the plurality of second component prediction areas;
[0342] A third component prediction block is determined or obtained according to the prediction results of the plurality of third component prediction areas.
[0343] Optionally, the prediction module A20 is further configured to perform at least one of the following:
[0344] Reconstructing the block based on the prediction residual obtained from the bitstream and the second component prediction block, and determining or generating a reconstruction of the block;
[0345] Reconstructing the block based on the prediction residual obtained from the bitstream and the third component prediction block, and determining or generating a reconstruction of the block;
[0346] Determining or generating a prediction residual based on the pixels of the block to be predicted and the second component prediction block;
[0347] A prediction residual is determined or generated according to the pixels of the block to be predicted and the third component prediction block.
[0348] Optionally, the prediction region includes a second component prediction region and / or a third component prediction region.
[0349] Optionally, the second component prediction area includes a rectangular area and / or a non-rectangular area;
[0350] The third component prediction area includes a rectangular area and / or a non-rectangular area;
[0351] Determining or obtaining a prediction model and / or model parameters corresponding to at least one prediction area based on reference pixels of the block to be predicted;
[0352] Determining or obtaining a cross-component prediction model and / or model parameters of the second component prediction region based on reference pixels of the second component prediction region and / or reference pixels of the third component prediction region;
[0353] The cross-component prediction model and / or model parameters of the third component prediction region are determined or obtained according to the reference pixels of the second component prediction region and / or the reference pixels of the third component prediction region.
[0354] Optionally, the reference pixels of the second component prediction area include neighboring pixels of the first component of the block to be predicted and neighboring pixels of the second component of the block to be predicted;
[0355] The reference pixels of the third component prediction area include neighboring pixels of the first component of the block to be predicted and neighboring pixels of the third component of the block to be predicted.
[0356] An embodiment of the present application further provides a processing device, including a memory and a processor, wherein a processing program is stored in the memory, and when the processing program is executed by the processor, the steps of the processing method in any of the above embodiments are implemented.
[0357] An embodiment of the present application further provides a storage medium having a processing program stored thereon. When the processing program is executed by a processor, the steps of the processing method in any of the above embodiments are implemented.
[0358] In the embodiments of the processing device, processing equipment and storage medium provided in this application, all technical features of any of the above-mentioned processing method embodiments may be included. The expansion and explanation content of the specification are basically the same as those of the embodiments of the above-mentioned methods and will not be repeated here.
[0359] An embodiment of the present application further provides a computer program product, which includes computer program code. When the computer program code runs on a computer, the computer executes the methods in the various possible implementation modes described above.
[0360] An embodiment of the present application also provides a chip, including a memory and a processor, wherein the memory is used to store computer programs, and the processor is used to call and run the computer programs from the memory, so that a device equipped with the chip executes the methods in the various possible implementation modes as described above.
[0361] It is understood that the above scenarios are merely examples and do not limit the application scenarios of the technical solutions provided in the embodiments of this application. The technical solutions of this application can also be applied to other scenarios. For example, those skilled in the art will appreciate that with the evolution of system architecture and the emergence of new business scenarios, the technical solutions provided in the embodiments of this application will also be applicable to similar technical problems.
[0362] The serial numbers of the above embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.
[0363] The steps in the method of the embodiment of the present application can be adjusted in order, combined and deleted according to actual needs.
[0364] The units in the device of the embodiment of the present application can be merged, divided and deleted according to actual needs.
[0365] In this application, the same or similar terminology, technical solutions and / or application scenario descriptions are generally only described in detail the first time they appear. When they appear again later, they are generally not repeated for the sake of brevity. When understanding the technical solutions and other contents of this application, for the same or similar terminology, technical solutions and / or application scenario descriptions that are not described in detail later, you can refer to the previous relevant detailed descriptions.
[0366] In this application, the description of each embodiment has its own focus. For parts that are not described or recorded in detail in a certain embodiment, please refer to the relevant description of other embodiments.
[0367] The various technical features of the technical solution of this application can be combined arbitrarily. In order to make the description concise, not all possible combinations of the various technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.
[0368] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus the necessary general hardware platform, and of course can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) as mentioned above, and includes a number of instructions for enabling a terminal device (which can be a mobile phone, computer, server, controlled terminal, or network device, etc.) to execute the method of each embodiment of the present application.
[0369] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When software is used for implementation, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function according to the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a storage medium or transmitted from one storage medium to another storage medium. For example, the computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) means. The storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more available media integrations. The available medium can be a magnetic medium (e.g., a floppy disk, a storage disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state storage disk Solid State Disk (SSD)).
[0370] The above are only preferred embodiments of the present application and do not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made using the contents of the present application specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present application.
Claims
1. A processing method, characterized in that, it includes steps: S10, dividing the block to be predicted into at least one prediction region; S20, performing prediction according to the prediction model and / or model parameters corresponding to the prediction region.
2. The processing method according to claim 1, characterized in that, the step S10 includes at least one of the following: dividing the block to be predicted into at least one prediction region according to the size information of the block to be predicted; dividing the block to be predicted into at least one prediction region according to the neighboring pixels of the block to be predicted; dividing the block to be predicted into at least one prediction region according to the division information of the block to be predicted; dividing the second component and / or the third component of the block to be predicted into at least one prediction region; dividing the block to be predicted into at least one prediction region according to the pixel gradient value of the neighboring block.
3. The processing method according to claim 1, characterized in that, the step S20 includes steps: S21, performing cross-component prediction on at least one prediction region according to the prediction model and / or model parameters corresponding to at least one prediction region, and determining or generating a prediction result; S22, determining or generating a prediction block according to the prediction results of the multiple prediction regions.
4. The processing method according to claim 3, characterized in that, the step S22 includes at least one of the following: determining or obtaining a second component prediction block according to the prediction results of multiple second component prediction regions; determining or obtaining a third component prediction block according to the prediction results of multiple third component prediction regions.
5. The processing method according to claim 4, characterized in that, it further includes at least one of the following: performing reconstruction according to the prediction residual obtained from the bitstream and the second component prediction block, and determining or generating the reconstruction of the block; performing reconstruction according to the prediction residual obtained from the bitstream and the third component prediction block, and determining or generating the reconstruction of the block; determining or generating a prediction residual according to the pixels of the block to be predicted and the second component prediction block; determining or generating a prediction residual according to the pixels of the block to be predicted and the third component prediction block.
6. The processing method according to any one of claims 1 to 5, characterized in that, the prediction region includes a second component prediction region and / or a third component prediction region.
7. The processing method according to claim 6, characterized in that, it further includes at least one of the following: the second component prediction region includes a rectangular region and / or a non-rectangular region; the third component prediction region includes a rectangular region and / or a non-rectangular region; determining or obtaining the prediction model and / or model parameters corresponding to at least one prediction region according to the reference pixels of the block to be predicted; determining or obtaining the cross-component prediction model and / or model parameters of the second component prediction region according to the reference pixels of the second component prediction region and / or the reference pixels of the third component prediction region; determining or obtaining the cross-component prediction model and / or model parameters of the third component prediction region according to the reference pixels of the second component prediction region and / or the reference pixels of the third component prediction region.
8. The processing method according to claim 7, characterized in that, it further includes at least one of the following: The reference pixels of the second component prediction region include the neighbor pixels of the first component of the block to be predicted and the neighbor pixels of the second component of the block to be predicted; The reference pixels of the third component prediction region include the neighbor pixels of the first component of the block to be predicted and the neighbor pixels of the third component of the block to be predicted.
9. A processing device, characterized in that it includes: a memory and a processor, wherein a processing program is stored on the memory, and when the processing program is executed by the processor, the steps of the processing method according to any one of claims 1 to 8 are implemented.
10. A storage medium, characterized in that a computer program is stored on the storage medium, and when the computer program is executed by a processor, the steps of the processing method according to any one of claims 1 to 8 are implemented.
Citation Information
Patent Citations
Processing method, processing equipment and storage medium
CN115955565A
Coding method, decoding method, electronic equipment and computer readable storage medium
CN116074537A
Image processing method, processing equipment and storage medium
CN116847088A
An apparatus, a method and a computer program for video coding and decoding
WO2023089230A2
Image processing method, intelligent terminal and storage medium
WO2023185351A1