Processing method, processing device, and storage medium

By employing a transform structure to transform the current block in a high-efficiency video coding standard protocol, the problem of insufficient transform structure in inter-frame prediction and intra-frame prediction is solved, thereby improving the bitstream compression efficiency.

CN121691725BActive Publication Date: 2026-06-26SHENZHEN TRANSSION HLDG CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SHENZHEN TRANSSION HLDG CO LTD
Filing Date
2026-02-12
Publication Date
2026-06-26

AI Technical Summary

Technical Problem

In high-efficiency video coding standard protocols, existing technologies fail to provide a non-independent transformation structure during the transformation process of inter-frame prediction and/or intra-frame prediction, resulting in large bit overhead in the bitstream and affecting the bitstream compression efficiency.

Method used

By transforming the current block according to the transform structure, including determining the transform method and transform structure, the bit overhead of the transform method for the current block in the bitstream is reduced. Transformation methods such as discrete cosine transform, transform skipping, sub-block transform, low-frequency non-separable quadratic transform, non-separable main transform, and multiple transform selection are provided to provide non-independent processing transform methods.

Benefits of technology

It improves the compression efficiency of the bitstream, reduces the bit overhead of the transformation method through non-independent processing, and enhances the encoding performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121691725B_ABST
    Figure CN121691725B_ABST
Patent Text Reader

Abstract

The application provides a processing method, a processing device and a storage medium. The processing method can be applied to the processing device and includes: transforming a current block according to at least one transform structure. According to the technical scheme, the transform structure can be used to provide a transform for non-independent processing (such as description, transmission or analysis) of the current block, so that the bit overhead of the transform mode of the transform of the current block in a code stream is reduced, and the code stream compression efficiency can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technology, specifically to a processing method, processing device, and storage medium. Background Technology

[0002] The existing high-efficiency video coding standard protocol (H.266 / VVC) proposes a video frame coding technique to improve coding performance without significantly increasing computational complexity. Specifically, when encoding and decoding video frames, the protocol divides each frame into different blocks and performs prediction, transformation and quantization processing before encoding and decoding.

[0003] In the process of conceiving and implementing this application, the inventors discovered at least the following problems: during the transformation process of inter-frame prediction and / or intra-frame prediction, the current block cannot be provided with non-independent processing (such as description, transmission or parsing) based on the transformation structure, resulting in large bit overhead related to the transformation process in the bitstream, which in turn affects the bitstream compression efficiency.

[0004] The preceding description is intended to provide general background information and does not necessarily constitute prior art. Summary of the Invention

[0005] To address the aforementioned technical problems, this application provides a processing method, processing device, and storage medium that can reduce the bit overhead of the transformation method for transforming the current block in the bitstream, thereby supporting improved bitstream compression efficiency.

[0006] This application provides a processing method applicable to a processing device, comprising the following steps:

[0007] S10, transform the current block according to at least one transformation structure.

[0008] Optionally, at least one transformation structure includes at least one transformation method.

[0009] Optionally, at least one transformation structure is determined or obtained based on at least one of the following:

[0010] The transformation structure of at least one sub-block of the current block;

[0011] List of candidate transformation structures;

[0012] The prediction pattern for the current block;

[0013] The current block size;

[0014] The size of at least one sub-block;

[0015] Bitstream.

[0016] Optionally, at least one of the transformation methods includes at least one of the following: discrete cosine transform, transform skip, sub-block transform, low-frequency non-separable quadratic transform, non-separable master transform, and multiple transform selection.

[0017] Optionally, the candidate transformation structure list includes: at least one candidate transformation structure and / or a list index of at least one candidate transformation structure.

[0018] Optionally, at least one candidate transformation structure is determined or obtained based on at least one of the following:

[0019] Default transformation structure;

[0020] The transformation structure of at least one of the following: the block above the current block, the block above the current block, the block to the left of the current block, the block to the left of the current block, the block to the upper left of the current block, and the block to the upper left of the current block;

[0021] Transformation structure of at least one of the following: the upper adjacent sub-block, the upper non-adjacent sub-block, the left adjacent sub-block, the left non-adjacent sub-block, the upper left adjacent sub-block, and the upper left non-adjacent sub-block;

[0022] Transformation structure of at least one of the following: cross-component block, same-component block, same-position block, and time-domain block of the current block;

[0023] Transformation structures of at least one sub-block across sub-blocks, same-component sub-blocks, co-position sub-blocks, and time-domain sub-blocks.

[0024] Optionally, at least one transformation structure is determined or obtained based on at least one candidate transformation structure in the candidate transformation structure list.

[0025] Optionally, at least one transformation structure is determined or obtained based on the transformation signaling in the code stream.

[0026] Optionally, the block size includes at least one of the following: width, height, aspect ratio, perimeter, and area.

[0027] Optionally, the list index of at least one candidate transformation structure is determined or obtained based on the sorting method of the candidate transformation structure list and / or the sorting position of at least one candidate transformation structure.

[0028] Optionally, the sorting position of at least one candidate transformation structure is determined or obtained based on at least one of the following:

[0029] The current block size;

[0030] The size of at least one sub-block;

[0031] The block size of at least one of the following: the block above the current block, the block above the current block, the block to the left of the current block, the block to the left of the current block, the block to the left of the current block, the block to the left of the current block, the block across components, the block in the same component, the block in the same position, and the block in the time domain.

[0032] The block size of at least one of the following: the upper adjacent sub-block, the upper non-adjacent sub-block, the left adjacent sub-block, the left non-adjacent sub-block, the upper left adjacent sub-block, the upper left non-adjacent sub-block, the cross-component sub-block, the same component sub-block, the same position sub-block, and the time domain sub-block;

[0033] The reuse frequency of at least one candidate transformation structure.

[0034] Optionally, the transformation signaling is determined or obtained based on at least one of the following:

[0035] Transformation structure identifier;

[0036] List index encoding result;

[0037] The number of candidate transformation structures;

[0038] Syntax elements for transformation methods;

[0039] Transform the kernel encoding result.

[0040] Optionally, at least one transformation structure identifier is determined or obtained based on the number of candidate transformation structures.

[0041] Optionally, at least one transformation structure identifier is determined or obtained based on the transformation cost.

[0042] Optionally, at least one list index encoding result includes at least one of the following: a first unary code encoding result, a first truncated unary code encoding result, and a first exponential Columbus encoding result.

[0043] Optionally, at least one transform kernel encoding result includes at least one of the following: a second unary code encoding result, a second truncated unary code encoding result, and a second exponential Columbus code result.

[0044] This application also provides a processing device, including: a memory and a processor, wherein the memory stores a processing program, and when the processing program is executed by the processor, it implements the steps of any of the processing methods described above.

[0045] This application also provides a storage medium storing a computer program that, when executed by a processor, implements the steps of any of the processing methods described above.

[0046] As described above, the processing method of this application can be applied to a processing device, including: transforming the current block according to at least one transformation structure. Through the technical solution of this application, a transformation method (such as description, transmission, or parsing) for the current block can be provided based on the transformation structure, thereby reducing the bit overhead of the transformation method for transforming the current block in the bitstream, and thus supporting improved bitstream compression efficiency. Attached Figure Description

[0047] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application. To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, those skilled in the art can obtain other drawings based on these drawings without any creative effort.

[0048] Figure 1 A schematic diagram of the hardware structure of a mobile terminal to implement the various embodiments of this application;

[0049] Figure 2 A communication network system architecture diagram provided for an embodiment of this application;

[0050] Figure 3 A schematic diagram of the hardware structure of a controller 140 provided in this application;

[0051] Figure 4 A schematic diagram of the hardware structure of a network node 150 provided in this application;

[0052] Figure 5 This is a flowchart illustrating the processing method according to the first embodiment;

[0053] Figure 6 This is a schematic diagram of the encoder's encoding process in the processing method shown in the first embodiment;

[0054] Figure 7 This is a schematic diagram of the decoding process of the decoder in the processing method shown in the first embodiment;

[0055] Figure 8 This is a schematic diagram of the spatial scene of adjacent sub-blocks shown according to the second embodiment;

[0056] Figure 9 This is a schematic diagram of a spatial scene of non-adjacent sub-blocks according to the second embodiment;

[0057] Figure 10 This is a schematic diagram of the spatial scene of the co-located sub-blocks shown in the second embodiment;

[0058] Figure 11 This is a flowchart illustrating the processing method according to the fourth embodiment. Figure 1 ;

[0059] Figure 12 This is a flowchart illustrating the processing method according to the fourth embodiment. Figure 2 ;

[0060] Figure 13This is a flowchart illustrating the processing method according to the fourth embodiment. Figure 3 ;

[0061] Figure 14 This is a flowchart illustrating the processing method according to the fourth embodiment. Figure 4 ;

[0062] Figure 15 This is a schematic diagram of the processing module of the processing device.

[0063] The realization of the objectives, functional features, and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. The accompanying drawings have illustrated specific embodiments of this application, which will be described in more detail below. These drawings and textual descriptions are not intended to limit the scope of the concept in any way, but rather to illustrate the concepts of this application to those skilled in the art through reference to specific embodiments. Detailed Implementation

[0064] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.

[0065] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, components, features, and elements with the same names in different embodiments of this application may have the same meaning or different meanings, the specific meaning of which must be determined by its interpretation in that specific embodiment or further in conjunction with the context of that specific embodiment.

[0066] It should be understood that although the terms first, second, third, etc., may be used herein to describe various information, such information should not be limited to these terms. These terms are used only to distinguish information of the same type from one another. For example, without departing from the scope of this document, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the word "if," as used herein, may be interpreted as "when," "when," or "in response to determination." Furthermore, as used herein, the singular forms "a," "an," and "the" are intended to also include the plural forms unless the context indicates otherwise. It should be further understood that the terms "comprising," "including," indicate the presence of the stated feature, step, operation, element, component, item, kind, and / or group, but do not exclude the presence, occurrence, or addition of one or more other features, steps, operations, elements, components, items, kinds, and / or groups. The terms "or," "and / or," "including at least one of the following," etc., as used in this application, may be interpreted as inclusive, or mean any one or any combination thereof. For example, "including at least one of the following: A, B, C" means "any one of the following: A; B; C; A and B; A and C; B and C; A and B and C." Similarly, "A, B, or C" or "A, B, and / or C" means "any one of the following: A; B; C; A and B; A and C; B and C; A and B and C." Exceptions to this definition only occur when the combination of elements, functions, steps, or operations is inherently mutually exclusive in some way.

[0067] It should be understood that although the steps in the flowcharts of this application's embodiments are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some of the steps in the figures may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be performed alternately or in turn with other steps or at least a portion of the sub-steps or stages of other steps.

[0068] Depending on the context, the words “if” or “suppose” as used here can be interpreted as “when” or “in response to determination” or “in response to detection.” Similarly, depending on the context, the phrases “if determination” or “if detection (of the stated condition or event)” can be interpreted as “when determination” or “in response to determination” or “when detection (of the stated condition or event)” or “in response to detection (of the stated condition or event).”

[0069] It should be noted that step designations such as S10 are used in this document for the purpose of more clearly and concisely describing the corresponding content, and do not constitute a substantial limitation on the order. In specific implementation, those skilled in the art may execute S20 first and then S10, etc., but these should all be within the protection scope of this application.

[0070] It should be understood that the specific embodiments described herein are merely illustrative of this application and are not intended to limit this application.

[0071] In the following description, the use of suffixes such as "module," "part," or "unit" to denote elements is solely for the purpose of illustrative purposes and has no specific meaning in itself. Therefore, "module," "part," or "unit" may be used interchangeably.

[0072] In this application, the terms "x", "N", and "M" appear. Unless otherwise specified, the value of a single lowercase letter x, a single uppercase letter N, or M is 0 or a positive integer.

[0073] In this application, the processing device can be a server (such as a local server or a cloud server) or a smart terminal. Optionally, the smart terminal can be implemented in various forms. For example, the smart terminal described in this application can include smart terminals such as mobile phones, tablets, laptops, handheld computers, personal digital assistants (PDAs), portable media players (PMPs), navigation devices, wearable devices, smart bracelets, pedometers, etc., as well as fixed terminals such as digital TVs and desktop computers.

[0074] The following description will use a mobile terminal as an example. Those skilled in the art will understand that, apart from elements specifically designed for mobile purposes, the construction according to the embodiments of this application can also be applied to fixed-type terminals.

[0075] Please see Figure 1 This is a schematic diagram of the hardware structure of a mobile terminal implementing various embodiments of this application. The mobile terminal 100 may include: an RF (Radio Frequency) unit 101, a WiFi module 102, an audio output unit 103, an A / V (Audio / Video) input unit 104, a sensor 105, a display unit 106, a user input unit 107, an interface unit 108, a memory 109, a processor 110, and a power supply 111, etc. Those skilled in the art will understand that... Figure 1The mobile terminal structure shown does not constitute a limitation on the mobile terminal. The mobile terminal may include more or fewer components than shown, or combine certain components, or have different component arrangements.

[0076] The following is combined with Figure 1 A detailed introduction to each component of the mobile terminal:

[0077] The radio frequency unit 101 can be used for receiving and transmitting signals during information transmission or calls. Specifically, it receives downlink information from the base station and processes it with the processor 110; additionally, it transmits uplink data to the base station. Typically, the radio frequency unit 101 includes, but is not limited to, an antenna, at least one amplifier, a transceiver, a coupler, a low-noise amplifier, and a duplexer. Furthermore, the radio frequency unit 101 can also communicate wirelessly with networks and other devices. The aforementioned wireless communications may use any communication standard or protocol, including but not limited to GSM (Global System of Mobile communication), GPRS (General Packet Radio Service), CDMA2000 (Code Division Multiple Access 2000), WCDMA (Wideband Code Division Multiple Access), TD-SCDMA (Time Division-Synchronous Code Division Multiple Access), FDD-LTE (Frequency Division Duplexing-Long Term Evolution), TDD-LTE (Time Division Duplexing-Long Term Evolution), 5G, and 6G.

[0078] WiFi is a short-range wireless transmission technology. Mobile terminals, through the WiFi module 102, can help users send and receive emails, browse web pages, and access streaming media, providing users with wireless broadband internet access. Although Figure 1 WiFi module 102 is shown, but it is understood that it is not a necessary component of a mobile terminal and can be omitted as needed without changing the nature of the invention.

[0079] The audio output unit 103 can convert audio data received by the radio frequency unit 101 or the WiFi module 102 or stored in the memory 109 into audio signals and output them as sound when the mobile terminal 100 is in call signal receiving mode, call mode, recording mode, voice recognition mode, broadcast receiving mode, etc. Furthermore, the audio output unit 103 can also provide audio output related to specific functions performed by the mobile terminal 100 (e.g., call signal receiving sound, message receiving sound, etc.). The audio output unit 103 may include a speaker, a buzzer, etc.

[0080] The A / V input unit 104 is used to receive audio or video signals. The A / V input unit 104 may include a graphics processing unit (GPU) 1041 and a microphone 1042. The GPU 1041 processes image data of still images or videos acquired by an image capture device (such as a camera) in video capture mode or image capture mode. The processed image frames can be displayed on the display unit 106. The image frames processed by the GPU 1041 can be stored in the memory 109 (or other storage media) or transmitted via the radio frequency unit 101 or the WiFi module 102. The microphone 1042 can receive sound (audio data) in operating modes such as telephone call mode, recording mode, and voice recognition mode, and can process such sound into audio data. The processed audio (voice) data can be converted into a format that can be transmitted to a mobile communication base station via the radio frequency unit 101 in telephone call mode. The microphone 1042 can implement various types of noise cancellation (or suppression) algorithms to eliminate (or suppress) noise or interference generated during the reception and transmission of audio signals.

[0081] The mobile terminal 100 also includes at least one sensor 105, such as a light sensor, a motion sensor, and other sensors. Optionally, the light sensor includes an ambient light sensor and a proximity sensor. Optionally, the ambient light sensor can adjust the brightness of the display panel 1061 according to the ambient light level, and the proximity sensor can turn off the display panel 1061 and / or backlight when the mobile terminal 100 is moved to the ear. As a type of motion sensor, the accelerometer sensor can detect the magnitude of acceleration in various directions (generally three axes), and can detect the magnitude and direction of gravity when stationary. It can be used for applications that recognize the phone's posture (such as landscape / portrait switching, related games, magnetometer posture calibration), vibration recognition related functions (such as pedometer, tapping), etc. Other sensors that may be configured in the phone, such as fingerprint sensors, pressure sensors, iris sensors, molecular sensors, gyroscopes, barometers, hygrometers, thermometers, and infrared sensors, will not be described in detail here.

[0082] The display unit 106 is used to display information input by the user or information provided to the user. The display unit 106 may include a display panel 1061, which may be configured in the form of a liquid crystal display (LCD), an organic light-emitting diode (OLED), or the like.

[0083] User input unit 107 can be used to receive input numerical or character information, and generate key signal inputs related to user settings and function control of the mobile terminal. Optionally, user input unit 107 may include touch panel 1071 and other input devices 1072. Touch panel 1071, also known as touch screen, can collect touch operations on or near the user (such as operations performed by the user using a finger, stylus, or any suitable object or accessory on or near touch panel 1071), and drive corresponding connection devices according to a pre-set program. Touch panel 1071 may include two parts: a touch detection device and a touch controller. Optionally, the touch detection device detects the user's touch position and the signal generated by the touch operation, and transmits the signal to the touch controller; the touch controller receives touch information from the touch detection device, converts it into touch point coordinates, sends it to processor 110, and can receive and execute commands sent by processor 110. In addition, touch panel 1071 can be implemented using various types such as resistive, capacitive, infrared, and surface acoustic wave. In addition to the touch panel 1071, the user input unit 107 may also include other input devices 1072. Optionally, other input devices 1072 may include, but are not limited to, one or more of the following: physical keyboard, function keys (such as volume control buttons, power buttons, etc.), trackball, mouse, joystick, etc., without being specifically limited here.

[0084] Optionally, the touch panel 1071 may cover the display panel 1061. When the touch panel 1071 detects a touch operation on or near it, it transmits the information to the processor 110 to determine the type of touch event. Subsequently, the processor 110 provides corresponding visual output on the display panel 1061 based on the type of touch event. Although in Figure 1 In this embodiment, the touch panel 1071 and the display panel 1061 are two independent components to realize the input and output functions of the mobile terminal. However, in some embodiments, the touch panel 1071 and the display panel 1061 can be integrated to realize the input and output functions of the mobile terminal. The specific implementation is not limited here.

[0085] Interface unit 108 serves as an interface through which at least one external device can connect to mobile terminal 100. For example, the external device may include a wired or wireless headset port, an external power supply (or battery charger) port, a wired or wireless data port, a memory card port, a port for connecting a device with an identification module, an audio input / output (I / O) port, a video I / O port, a headphone port, and so on. Interface unit 108 may be used to receive input (e.g., data, power, etc.) from the external device and transmit the received input to one or more elements within mobile terminal 100, or it may be used to transmit data between mobile terminal 100 and the external device.

[0086] The memory 109 can be used to store software programs and various data. The memory 109 may primarily include a program storage area and a data storage area. Optionally, the program storage area may store the operating system, applications required for at least one function (such as sound playback, image playback, etc.), etc.; the data storage area may store data created based on the use of the mobile phone (such as audio data, phonebook, etc.). Furthermore, the memory 109 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device.

[0087] The processor 110 is the control center of the mobile terminal. It connects various parts of the mobile terminal via various interfaces and lines. By running or executing software programs and / or modules stored in the memory 109, and by calling data stored in the memory 109, it performs various functions and processes data of the mobile terminal, thereby providing overall monitoring of the mobile terminal. The processor 110 may include one or more processing units; preferably, the processor 110 may integrate an application processor and a modem processor. Optionally, the application processor mainly handles the operating system, user interface, and applications, while the modem processor mainly handles wireless communication. It is understood that the modem processor may not be integrated into the processor 110.

[0088] The mobile terminal 100 may also include a power supply 111 (such as a battery) that supplies power to various components. Preferably, the power supply 111 can be logically connected to the processor 110 through a power management system, thereby enabling functions such as charging, discharging, and power consumption management through the power management system.

[0089] although Figure 1 As not shown, the mobile terminal 100 may also include a Bluetooth module, etc., which will not be described in detail here.

[0090] To facilitate understanding of the embodiments of this application, the communication network system on which the mobile terminal of this application is based is described below.

[0091] Please see Figure 2 , Figure 2 This application provides a communication network system architecture diagram. The communication network system is an LTE system based on the universal mobile communication technology. The LTE system includes a UE (User Equipment) 201, an E-UTRAN (Evolved UMTS Terrestrial Radio Access Network) 202, an EPC (Evolved Packet Core) 203, and the operator's IP services 204, which are connected in sequence.

[0092] Optionally, UE201 can be the aforementioned mobile terminal 100, which will not be described in detail here.

[0093] E-UTRAN202 includes eNodeB2021 and other eNodeB2022s. Optionally, eNodeB2021 can connect to other eNodeB2022s via backhaul (e.g., X2 interface). eNodeB2021 connects to EPC203 and can provide UE201 with access to EPC203.

[0094] EPC203 may include an MME (Mobility Management Entity) 2031, an HSS (Home Subscriber Server) 2032, other MMEs 2033, an SGW (Serving Gateway) 2034, a PGW (Packet Data Network Gateway) 2035, and a PCRF (Policy and Charging Rules Function) 2036, etc. Optionally, MME2031 is the control node that handles signaling between UE201 and EPC203, providing bearer and connection management. HSS2032 is used to provide registers to manage functions such as the Home Location Register (not shown in the figure) and stores user-specific information such as service characteristics and data rates. All user data can be sent through SGW2034. PGW2035 can provide UE 201 IP address allocation and other functions. PCRF2036 is the policy and charging control decision point for service data flow and IP bearer resources. It selects and provides available policy and charging control decisions for the policy and charging enforcement function unit (not shown in the figure).

[0095] IP services 204 may include the Internet, intranet, IMS (IP Multimedia Subsystem), or other IP services.

[0096] Although the above description uses the LTE system as an example, those skilled in the art should know that this application is not only applicable to the LTE system, but also to other wireless communication systems, such as GSM, CDMA2000, WCDMA, TD-SCDMA, 5G and future new network systems (such as 6G), etc., without limitation.

[0097] Figure 3 This is a schematic diagram of the hardware structure of a controller 140 provided in this application. The controller 140 includes a memory 1401 and a processor 1402. The memory 1401 is used to store program instructions, and the processor 1402 is used to call the program instructions in the memory 1401 to execute the steps performed by the controller in the first embodiment of the above method. The implementation principle and beneficial effects are similar, and will not be described again here.

[0098] Optionally, the controller further includes a communication interface 1403, which can be connected to the processor 1402 via a bus 1404. The processor 1402 can control the communication interface 1403 to implement the receiving and sending functions of the controller 140.

[0099] Figure 4 This application provides a schematic diagram of the hardware structure of a network node 150. The network node 150 includes a memory 1501 and a processor 1502. The memory 1501 is used to store program instructions, and the processor 1502 is used to call the program instructions in the memory 1501 to execute the steps performed by the first node in the first embodiment of the above method. The implementation principle and beneficial effects are similar, and will not be described again here.

[0100] Optionally, the controller further includes a communication interface 1503, which can be connected to the processor 1502 via a bus 1504. The processor 1502 can control the communication interface 1503 to implement the receiving and sending functions of the network node 150.

[0101] The integrated modules described above, implemented as software functional modules, can be stored in a computer-readable storage medium. These software functional modules, stored in a storage medium, include several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute some steps of the methods of the various embodiments of this application.

[0102] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product. A computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the flow or function according to the embodiments of this application is generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a storage medium or transmitted from one storage medium to another. For example, computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state disk, SSD), etc.

[0103] Based on the above-described mobile terminal hardware structure and communication network system, various embodiments of this application are proposed.

[0104] First Embodiment

[0105] Reference Figure 5 , Figure 5 This is a flowchart illustrating the processing method according to the first embodiment. The processing method of this application embodiment can be applied to a processing device and includes the following step S10:

[0106] Step S10: Transform the current block according to at least one transformation structure;

[0107] In this embodiment, the processing device can be a smart terminal, such as a mobile phone or a computer, and / or the processing device can be a server, such as a local server or a cloud server. In this embodiment and this application, the processing device is mainly described as a smart terminal.

[0108] Optionally, the technical solution of this embodiment can be applied to fields such as image encoding and decoding, video encoding and decoding, hardware video encoding and decoding, dedicated circuit video encoding and decoding, and / or real-time video encoding and decoding.

[0109] Optionally, the prediction mode used by the current block in this application embodiment can be an intra-frame prediction mode and / or an inter-frame prediction mode.

[0110] Optionally, at least one transformation structure includes at least one transformation method.

[0111] Optionally, the transformation structure can be one transformation method and / or a combination of multiple transformation methods for performing residual transformation on the current block. It is used to uniformly describe the combination of multiple transformation methods to be performed on the current block during the residual transformation process, in order to replace the method of processing each transformation method independently (such as describing, transmitting or parsing).

[0112] Optionally, the processing device can be either an encoder or a decoder. When the processing device is an encoder and / or a decoder, it can perform a transformation of the current block according to at least one transformation structure.

[0113] Optionally, when the processing device is an encoder, the encoder can obtain video image data from the video source, segment each frame of the video image data to obtain at least one image block, and the image block undergoing transformation processing in the at least one image block is the current block.

[0114] Optionally, when the processing device is a decoder, the decoder can decode the data in the bitstream to obtain at least one image block, and the image block undergoing transformation processing in the at least one image block is the current block.

[0115] Optionally, the transformation process can be a mathematical transformation operation performed on the prediction residual of the current block during the image reconstruction stage to convert the spatial domain residual data to the frequency domain to generate corresponding transformation coefficients, and / or convert the transformation coefficients from the frequency domain to the spatial domain residual data, thereby reducing spatial redundancy of the data, facilitating subsequent quantization and entropy coding, and providing a basis for reconstructing the image.

[0116] Optionally, the current block can be a current coding block or a current decoding block. The current coding block can be the image block that is currently being encoded and processed in the encoder, while the current decoding block can be the image block that is currently being reconstructed in the decoder.

[0117] Optionally, the current block can be a Coding Tree Unit (CTU), a Macroblock, and / or a Coding Unit (CU).

[0118] Optionally, at least one of the transformation methods includes at least one of the following: Discrete Cosine Transform (DCT), Transform Skip (TS), Sub-Block Transform (SBT), Low-Frequency Non-Separable Secondary Transform (LFNST), Non-Separable Primary Transform (NSPT), and Multiple Transform Selection (MTS).

[0119] Optionally, DCT can be a second-class discrete cosine transform (DCT-2), an orthogonal transform method used to transform the prediction residuals in the spatial domain to the frequency domain. By concentrating energy on a few low-frequency coefficients, it facilitates subsequent quantization and entropy coding, thereby achieving efficient data compression.

[0120] Optionally, TS skips the transformation process of the prediction residual and performs quantization and entropy encoding on the prediction residual in the spatial domain. Transform skipping can be used for image patches with sharp edges and / or non-stationary characteristics, such as screen content. By performing quantization and entropy encoding, energy diffusion caused by the transformation can be avoided, and the original image features can be better preserved.

[0121] Optionally, SBT is a transformation method that combines partitioning with bound transform kernels. Its core idea is not to transform the entire CU, but to divide the CU into at least two transform units (TUs), and select only one TU for transformation, while the prediction residual of the other TU is forced to zero, which can represent effective information with fewer bits.

[0122] Optionally, LFNST is a secondary processing transformation performed after the main transformation (for example, in the VVC protocol, DCT-2 is a secondary processing transformation performed after the main transformation). This transformation method does not operate on all coefficients after the main transformation, but only applies a two-dimensional non-separable transformation kernel to further transform the low-frequency coefficient region. This is used to enhance the energy concentration in the low-frequency coefficient region and / or mine the correlation between coefficients, thereby improving compression efficiency.

[0123] Optionally, NSPT is a form of non-separable master transform that operates on prediction residuals. This transform method uses a two-dimensional non-separable transform kernel instead of a traditional separable master transform kernel to process image patches at once, so as to better adapt to the statistical characteristics of the prediction residuals of image patches in various directions.

[0124] Optionally, MTS is a mechanism that provides multiple candidate transform kernel combinations for prediction residuals. It selects a set of row transform kernels and column transform kernels from a predefined set of candidate transform kernels to adapt to the frequency and directional characteristics of different prediction residuals.

[0125] Optionally, the processing method of this application embodiment, when applied to the encoder and / or decoder, can replace the traditional transformation method, and / or be a sub-method of the traditional transformation method, and / or be combined with the traditional transformation method. For example, after transforming the current block according to at least one traditional transformation method, the current block can be subsequently transformed according to a transformation structure that does not include the transformation method, or after transforming the current block according to at least one transformation structure, the current block can be transformed according to a traditional transformation method that is different from any of the transformation methods in the transformation structure, etc. This application embodiment does not limit this.

[0126] Reference Figure 6 When the processing device is an encoder on the encoding side, the encoder can receive video data input from a video source, such as receiving video images from the video source, determining the image to be predicted in the video images, dividing the image to be predicted into at least one image block, the image block including luma blocks and chroma blocks, and using the temporal and / or spatial correlation between video images, performing prediction processing on at least one image block, including intra-frame prediction processing and / or inter-frame prediction processing. Intra-frame prediction processing and / or inter-frame prediction processing each include at least one prediction mode. For the above prediction modes, the encoder uses, for example, rate-distortion cost to determine the prediction mode finally adopted by at least one image block. For example, it calculates the rate-distortion cost corresponding to at least one prediction mode and / or the rate-distortion cost of combining several prediction modes to determine the minimum rate-distortion cost from at least one rate-distortion cost. The prediction mode corresponding to the minimum rate-distortion cost and / or the combination of prediction modes is the prediction mode finally adopted by the image block, and the prediction mode includes intra-frame prediction mode and / or inter-frame prediction mode.

[0127] Optionally, if the prediction block corresponding to the image block is determined or obtained, the pixel value of the pixel sample in the original image block corresponding to the image block is subtracted from the prediction value of the corresponding pixel sample in the prediction block to obtain the residual value of the pixel sample and / or the residual block corresponding to the original image block. The residual block can be transformed and / or quantized and encoded by an entropy encoder to form an encoded bit stream (i.e., code stream).

[0128] Optionally, the transformed and / or quantized residual block can be added to the prediction block determined or obtained through the prediction mode to obtain the reconstructed block. The reconstructed block can also be subjected to loop filtering to reduce distortion.

[0129] Optionally, the loop filtering process includes: deblocking filtering, adaptive sampling offset and / or adaptive loop filtering, and after the loop filtering process is performed, the reconstructed block after the loop filtering process is stored according to the coded image buffer.

[0130] Optionally, the encoded bitstream may include prediction parameters and related auxiliary information corresponding to a determined prediction mode, and / or filtering parameters and related auxiliary information corresponding to a determined filtering mode, and the aforementioned prediction parameters and / or filtering parameters are packaged into the encoded bitstream after entropy encoding.

[0131] Optionally, the transformation process may include transforming the current block according to at least one transformation structure as proposed in the embodiments of this application.

[0132] Reference Figure 7 When the processing device is a decoder on the decoding side, after receiving the encoded bit stream, the decoder's entropy decoding unit will parse and / or decode the encoded bit stream to obtain the transform coefficients. The decoder's dequantization unit and inverse transform unit will perform dequantization and inverse transform processing on the transform coefficients to obtain the residual block.

[0133] Optionally, the decoding unit of the decoder parses and decodes the encoded bitstream to obtain prediction parameters and related auxiliary information. The prediction processing unit of the decoder uses the prediction parameters to perform prediction processing to determine the prediction block corresponding to the residual block. The prediction processing includes intra-frame prediction processing and / or inter-frame prediction processing, which includes a combination of one or more prediction modes.

[0134] Optionally, the prediction result of the image block (i.e., the prediction block) is determined or obtained according to the prediction mode indicated by the auxiliary information, and the determined or obtained residual block and prediction block (including the predicted luminance block and / or the predicted chrominance block) are added together to obtain the reconstructed block.

[0135] Optionally, the decoder can perform loop filtering on the reconstructed blocks according to the filtering mode indicated by the auxiliary information to reduce distortion and improve video quality. The reconstructed blocks after loop filtering are further combined into a decoded image and stored in the decoded image buffer or output as a decoded video signal.

[0136] Optionally, the loop filtering process includes: deblocking filtering, adaptive sampling offset, and / or adaptive loop filtering.

[0137] Optionally, the transformation process may include transforming the current block according to at least one transformation structure as proposed in the embodiments of this application.

[0138] In this embodiment, by transforming the current block through at least one transformation structure, a transformation method (such as description, transmission, or parsing) can be provided for the current block based on the transformation structure, thereby reducing the bit overhead of the transformation method for transforming the current block in the bitstream, and thus supporting the improvement of bitstream compression efficiency.

[0139] Second Embodiment

[0140] Based on the first embodiment described above, a second embodiment is proposed.

[0141] In this embodiment, at least one transformation structure is determined or obtained according to at least one of the following methods a1 to a6:

[0142] Method a1, the transformation structure of at least one sub-block of the current block;

[0143] Optionally, the transformation structure of at least one sub-block of the current block can be used as the transformation structure for transforming the current block.

[0144] Optionally, the sub-block can be a sub-block within the current block, or a sub-block with pre-defined dimensions, such as an 8x8 sub-block, and the size parameters of the sub-block (such as length, width, height, perimeter, and area) are smaller than the size parameters of the current block.

[0145] Optionally, the sub-blocks of the current block are not separate image blocks, but rather pre-determined image block regions within the current block. The current block may include one or more sub-blocks. That is, it can be inferred in advance that the current block can be divided into at least one sub-block, but no division processing is performed at this time, and it is used as a sub-block of the current block. For example, a 16x16 image block may include four 8x8 sub-blocks.

[0146] Optionally, the sub-block can be any one of at least two TUs obtained by dividing the current block according to the SBT transformation method;

[0147] Optionally, the transformation structure of a sub-block can be a transformation method and / or a combination of multiple transformation methods for performing residual transformation on the sub-block. This is used to uniformly describe the combination of multiple transformation methods to be performed on the sub-block during the residual transformation process, instead of processing each transformation method independently (such as describing, transmitting, or parsing).

[0148] Optionally, the transformation process can be a mathematical transformation operation performed on the prediction residuals of sub-blocks during the image reconstruction stage to convert the spatial domain residual data to the frequency domain to generate corresponding transformation coefficients, and / or convert the transformation coefficients from the frequency domain to the spatial domain residual data, thereby reducing spatial redundancy of data, facilitating subsequent quantization and entropy coding, and providing a basis for reconstructing the image.

[0149] Optionally, the transformation structure of at least one sub-block can be a combination of one or more transformation methods pre-defined for the sub-block, and / or it can be a transformation structure that reuses one of the following: neighboring sub-blocks, non-neighboring sub-blocks, cross-component sub-blocks, same-component sub-blocks, co-position sub-blocks, and time-domain sub-blocks of the sub-block.

[0150] Optionally, the neighboring sub-block can be a sub-block adjacent to the current sub-block (such as the upper adjacent sub-block, the left adjacent sub-block, or the upper left adjacent sub-block), and / or can be a sub-block that has already been predicted or reconstructed.

[0151] Optionally, a non-neighboring sub-block can be a sub-block that is not adjacent to the current sub-block (such as a non-adjacent sub-block above, a non-adjacent sub-block to the left, or a non-adjacent sub-block to the upper left), and / or can be a sub-block that has already been predicted or reconstructed.

[0152] Optionally, the cross-component sub-block can be an image block that is in a different component from at least one sub-block. For example, if at least one sub-block to be transformed is an image block of the Y component, then the cross-component sub-block can be an image block of the U component and / or V component.

[0153] Optionally, if at least one sub-block is an image block of the U component, then the cross-component sub-block can be an image block of the Y component and / or the V component.

[0154] Optionally, if at least one sub-block is an image block of the V component, then the cross-component sub-block can be an image block of the Y component and / or the U component.

[0155] Optionally, the sub-block of the same component can be an image block that is in the same component as at least one sub-block. For example, if at least one sub-block to be transformed is an image block of the Y component, then the sub-block of the same component can be an image block of the Y component.

[0156] Optionally, if at least one sub-block is an image block of the U component, then sub-blocks of the same component can be image blocks of the U component.

[0157] Optionally, if at least one sub-block is an image block of the V component, then sub-blocks of the same component can be image blocks of the V component.

[0158] Optionally, a co-positional sub-block can be an image block in a co-positional image that has the same position and size as the current sub-block and / or is in the same co-positional block. Optionally, a co-positional image can be the image in a reference image that is closest in time to the current image.

[0159] Optionally, the temporal sub-block can be a sub-block that is distinguished in the time domain, such as a sub-block of the image block in the previous frame. For example, if there is video data containing three frames of images, the first frame is played in the first second, the second frame is played in the second second, and the third frame is played in the third second, if the image block predicted at the current moment (such as the current block) is the image block after the second frame is divided, then the temporal block can be determined or obtained as a sub-block of the image block corresponding to it in the first frame.

[0160] Optionally, the transformation structure of at least one sub-block of the current block can be used as the transformation structure for transforming the current block.

[0161] In this approach, the transformation decision process is simplified by using the transformation structure of at least one sub-block of the current block as the transformation structure for transforming the current block. The transformation structure configuration applicable to the whole (current block) is derived based on the transformation structure of the local (sub-block), thereby reducing the computational complexity of determining or obtaining the transformation structure and improving the encoding and decoding performance.

[0162] Method a2, candidate transformation structure list;

[0163] Optionally, the candidate transformation structure list may be a list built in advance and / or in real time to provide at least one optional candidate transformation structure for determining or obtaining at least one transformation structure.

[0164] Optionally, the candidate transformation structure list includes: at least one candidate transformation structure and / or a list index of at least one candidate transformation structure.

[0165] Optionally, the candidate transformation structure can be a transformation structure recorded in the candidate transformation structure list.

[0166] Optionally, the list index can be an index identifier that corresponds one-to-one with the candidate transformation structure in the candidate transformation structure list. The list index can be a number and / or a character. One list index corresponds to one candidate transformation structure and is used to indicate the position and / or order of the candidate transformation structure in the candidate transformation structure list (for example, list index "1" indicates the first candidate transformation structure recorded in the candidate transformation structure list, and list index "2" indicates the second candidate transformation structure recorded in the candidate transformation structure list).

[0167] Optionally, at least one candidate transformation structure can be selected from the candidate transformation structure list as the transformation structure for the current block, so as to transform the current block according to the selected transformation structure.

[0168] Optionally, the selection criteria for a candidate transformation structure in the candidate transformation structure list can be rate-distortion cost and / or list index.

[0169] Optionally, when the processing device is an encoder, if the number of candidate transformation structures recorded in the candidate transformation structure list is one, the encoder may select the recorded candidate transformation structure as the transformation structure for transforming the current block.

[0170] Optionally, when the processing device is an encoder, if the number of candidate transformation structures recorded in the candidate transformation structure list is greater than one, the encoder can perform rate-distortion optimization on each candidate transformation structure recorded in the candidate transformation structure list, so as to select the candidate transformation structure with the lowest cost based on the cost of rate-distortion as the transformation structure to transform the current block.

[0171] Optionally, when the processing device is a decoder, the decoder can determine or obtain the candidate transform structure recorded in the candidate transform structure list corresponding to the list index according to the list index in the bitstream, so as to use the candidate transform structure corresponding to the list index as the transform structure to transform the current block.

[0172] Optionally, the candidate transform structure list can be constructed from the encoder and / or decoder.

[0173] Optionally, the list of candidate transform structures built at the encoder and decoder is the same.

[0174] Optionally, when there are at least one identical candidate transformation structure in the candidate transformation structure list, the at least one identical candidate transformation structure can be deduplicated, and one candidate transformation structure can be retained among the at least one identical candidate transformation structure.

[0175] In this approach, by using a candidate transform structure list as the source of transform structures for transforming the current block, multiple selectable transform structures (e.g., rate-distortion optimization) can be provided for the current block, thereby supporting the improvement of the transform adaptability of the determined or obtained transform structures, and thus supporting the improvement of encoding and / or decoding quality in the video encoding and / or decoding process.

[0176] Optionally, at least one candidate transformation structure list is determined or obtained according to at least one of the following methods a21 to a25:

[0177] Method a21, default transformation structure;

[0178] Optionally, the default transform structure can be a pre-set transform structure, such as a typical transform structure pre-set by the encoder and / or decoder.

[0179] Optionally, the order of the transformation modes in the default transformation structure can be as follows: SBT off, DCT-2, TS, SBT on + IFNST, NSPT, DCT-2 + IFNST.

[0180] Optionally, at least one default transformation structure can be included as a candidate transformation structure in a list of at least one candidate transformation structures.

[0181] Method a22, a transformation structure of at least one of the following: the block above the current block, the block above the current block, the block to the left of the current block, the block to the left of the current block, the block to the upper left of the current block, and the block to the upper left of the current block;

[0182] Optionally, the adjacent block above can be an image block that is in the same image frame as the current block and is located above and adjacent to the current block.

[0183] Optionally, the non-adjacent block above can be an image block that is in the same image frame as the current block and is located above the current block but not adjacent to it. For example, when the current block is a CU, the non-adjacent block above is another CU that is offset by x (an integer greater than 1) CU heights above the current block.

[0184] Optionally, the left adjacent block can be an image block that is in the same image frame as the current block and is located to the left of the current block and adjacent to the current block.

[0185] Optionally, the non-adjacent block to the left can be an image block that is in the same image frame as the current block and is located to the left of the current block but is not adjacent to the current block. For example, when the current block is a CU, the non-adjacent block to the left is another CU that is offset to the left of the current block by x (an integer greater than 1) CU heights.

[0186] Optionally, the upper left adjacent block can be an image block that is in the same image frame as the current block and is located to the upper left of the current block and adjacent to the current block.

[0187] Optionally, the upper left non-adjacent block can be an image block that is in the same image frame as the current block and is located to the upper left of the current block but is not adjacent to the current block. For example, when the current block is a CU, the upper left non-adjacent block is another CU that is offset by x (an integer greater than 1) CU heights above and to the left of the current block.

[0188] Optionally, the transformation structure of at least one of the above adjacent block, above non-adjacent block, left adjacent block, left non-adjacent block, upper left adjacent block, and upper left non-adjacent block of the current block is used as a candidate transformation structure in the candidate transformation structure list to realize spatial reuse of transformation structure at the scale of the current block.

[0189] Optionally, when spatially reusing the transformation structures of the upper adjacent block, upper non-adjacent block, left adjacent block, left non-adjacent block, upper left adjacent block and / or upper left non-adjacent block of the current block, if there are multiple identical transformation structures in the reused and / or to be reused transformation structures, the multiple identical transformation structures can be deduplicated, and one of the multiple identical transformation structures can be retained as a candidate transformation structure.

[0190] Method a23, a transformation structure of at least one of the following: the upper adjacent sub-block, the upper non-adjacent sub-block, the left adjacent sub-block, the left non-adjacent sub-block, the upper left adjacent sub-block, and the upper left non-adjacent sub-block;

[0191] Optionally, the adjacent sub-block above can be another sub-block that is in the same image frame as the current block and is located above and adjacent to the sub-block of the current block.

[0192] Optionally, the left adjacent sub-block can be another sub-block that is in the same image frame as the current block and is located to the left of the sub-block of the current block and adjacent to that sub-block.

[0193] Optionally, the upper left adjacent sub-block can be another sub-block that is in the same image frame as the current block and is located to the upper left of the sub-block of the current block and adjacent to the current sub-block.

[0194] Reference Figure 8 The sub-block above the sub-block includes block A1, the sub-block to the left of the sub-block includes block A2, the sub-block to the upper left of the sub-block includes block A5, block A3 is the sub-block to the upper right of the sub-block, and block A4 is the sub-block to the lower left of the sub-block. Blocks A3 and A4 are adjacent sub-blocks that have not been predicted and / or encoded / decoded, and do not have the reference value for transform structure reuse. Therefore, the candidate transform structure list of the sub-block is determined or obtained only based on the transform structure of at least one of blocks A1, A2 and / or A5.

[0195] Optionally, the non-adjacent sub-block above can be another sub-block that is in the same image frame as the current block and is located above the sub-block of the current block but is not adjacent to it. For example, when the sub-block of the current block is a TU, the non-adjacent sub-block above is another TU that is offset by x (an integer greater than 1) TU heights above the sub-block.

[0196] Optionally, the non-adjacent sub-block on the left can be another sub-block that is in the same image frame as the current block and is located to the left of the sub-block of the current block but is not adjacent to that sub-block. For example, when the sub-block of the current block is a TU, the non-adjacent sub-block on the left is another TU that is offset to the left of the sub-block by x (an integer greater than 1) TU heights.

[0197] Optionally, the upper left non-adjacent sub-block can be another sub-block that is in the same image frame as the current block and is located to the upper left of the sub-block of the current block but is not adjacent to the current sub-block. For example, when the sub-block of the current block is a TU, the upper left non-adjacent sub-block is another TU that is offset by x (an integer greater than 1) TU heights above and to the left of the current sub-block.

[0198] Reference Figure 9The non-adjacent sub-blocks above the sub-block include B1, which is offset by x TU blocks in the height H direction relative to the sub-block TU. The non-adjacent sub-blocks to the left of the sub-block include B2, which is offset by x TU blocks in the width W direction relative to the TU. The non-adjacent sub-blocks to the upper left of the sub-block include B5, which is offset by x TU blocks in both the height H and width W directions relative to the TU. Block B3 is the non-adjacent sub-block to the upper right of the sub-block, and block B4 is the non-adjacent sub-block to the lower left of the sub-block. Blocks B3 and B4 are non-adjacent sub-blocks that have not been predicted and / or encoded / decoded, and do not have the reference value for transform structure reuse. Therefore, the candidate transform structure list of the sub-block is determined or obtained only based on the transform structure of at least one of blocks B1, B2, and / or B5.

[0199] Optionally, the transformation structure of at least one of the above adjacent sub-block, above non-adjacent sub-block, left adjacent sub-block, left non-adjacent sub-block, left upper adjacent sub-block, and left upper non-adjacent sub-block of at least one sub-block of the current block is used as a candidate transformation structure in the candidate transformation structure list, thereby realizing the spatial reuse of transformation structures at the scale of sub-blocks.

[0200] Optionally, when multiplexing the transformation structures of the upper adjacent sub-block, upper non-adjacent sub-block, left adjacent sub-block, left non-adjacent sub-block, upper left adjacent sub-block and / or upper left non-adjacent sub-block of a spatial multiplexing sub-block, if there are multiple identical transformation structures in the multiple identical transformation structures, the multiple identical transformation structures can be deduplicated, and one of the multiple identical transformation structures can be retained as a candidate transformation structure.

[0201] Method a24, a transformation structure of at least one of the following: cross-component block, same-component block, same-position block and time-domain block of the current block;

[0202] Optionally, the cross-component block can be an image block in a different component than the current block. For example, when the current block is an image block of the Y component, the cross-component block of the current block can be an image block of the U component and / or the V component.

[0203] Optionally, when the current block is an image block of the U component, the cross-component block of the current block can be an image block of the Y component and / or the V component.

[0204] Optionally, when the current block is an image block of the V component, the cross-component block of the current block can be an image block of the Y component and / or the U component.

[0205] Optionally, a component block can be an image sub-block that is in the same component as the current block. For example, if the current block is an image block of the Y component, then a component block of the current block can be an image block of another Y component.

[0206] Optionally, when the current block is an image block of the U component, the block of the same component of the current block can be an image block of another U component.

[0207] Optionally, when the current block is an image block of the V component, the same component block of the current block can be an image block of another V component.

[0208] Optionally, the co-location block can be an image block in the co-location image that has the same position and size as the current block. Optionally, the co-location image can be the image in the reference image that is closest to the current image in time.

[0209] Optionally, the temporal block can be a block that is distinguished in the time domain, such as an image block in the previous frame. For example, if there is video data containing three frames of images, the first frame is played in the first second, the second frame is played in the second second, and the third frame is played in the third second, if the image block predicted at the current moment (such as the current block) is an image block after the second frame is divided, then the temporal block can be determined or obtained as the image block corresponding to it in the first frame.

[0210] Optionally, the transform structure of at least one of the cross-component blocks, same-component blocks, same-position blocks and time-domain blocks of the current block is selected as a candidate transform structure in the candidate transform structure list, thereby realizing component multiplexing and / or time multiplexing of the transform structure at the scale of the current block.

[0211] Optionally, when multiplexing and / or time multiplexing the transformation structure of the current block across component blocks, same component blocks, same position blocks and time domain blocks, if there are multiple identical transformation structures in the multiplexed and / or to be multiplexed transformation structures, the multiple identical transformation structures can be deduplicated, and one of the multiple identical transformation structures can be retained as a candidate transformation structure.

[0212] Method a25, a transformation structure of at least one sub-block across component sub-blocks, same component sub-blocks, co-position sub-blocks and time-domain sub-blocks.

[0213] Optionally, the cross-component sub-block can be an image block that is in a different component from at least one sub-block. For example, if at least one sub-block to be transformed is an image block of the Y component, then the cross-component sub-block can be an image block of the U component and / or V component.

[0214] Optionally, if at least one sub-block is an image block of the U component, then the cross-component sub-block can be an image block of the Y component and / or the V component.

[0215] Optionally, if at least one sub-block is an image block of the V component, then the cross-component sub-block can be an image block of the Y component and / or the U component.

[0216] Optionally, the sub-block of the same component can be an image block that is in the same component as at least one sub-block. For example, if at least one sub-block to be transformed is an image block of the Y component, then the sub-block of the same component can be an image block of the Y component.

[0217] Optionally, if at least one sub-block is an image block of the U component, then sub-blocks of the same component can be image blocks of the U component.

[0218] Optionally, if at least one sub-block is an image block of the V component, then sub-blocks of the same component can be image blocks of the V component.

[0219] Optionally, a co-location sub-block can be an image block in a co-location image that has the same position and size as the current sub-block and / or is in the same co-location block.

[0220] Alternatively, the co-position image can be the reference image that is temporally closest to the current image.

[0221] Reference Figure 10 Among the co-position blocks in the co-position image, there is at least one of the following: block C1, which has the same position and size as the current sub-block; block C2, which is located at the top left corner of the same co-position block; block C3, which is located at the top right corner of the same co-position block; block C4, which is located at the bottom left corner of the same co-position block; and block C5, which is located at the bottom right corner of the same co-position block.

[0222] Optionally, the temporal sub-block can be a sub-block that is distinguished in the time domain, such as a sub-block of the image block in the previous frame. For example, if there is video data containing three frames of images, the first frame is played in the first second, the second frame is played in the second second, and the third frame is played in the third second, if the image block predicted at the current moment (such as the current block) is the image block after the second frame is divided, then the temporal block can be determined or obtained as a sub-block of the image block corresponding to it in the first frame.

[0223] Optionally, the transformation structure of at least one of the cross-component sub-blocks, same-component sub-blocks, co-position sub-blocks and time-domain sub-blocks of at least one sub-block of the current block is selected as a candidate transformation structure in the candidate transformation structure list, thereby realizing component multiplexing and / or time multiplexing of the transformation structure at the sub-block scale.

[0224] Optionally, when transforming the transformation structure of component multiplexing and / or time multiplexing sub-blocks across component sub-blocks, same component sub-blocks, co-position sub-blocks, and time-domain sub-blocks, if there are multiple identical transformation structures in the multiple multiplexed and / or to be multiplexed transformation structures, the multiple identical transformation structures can be deduplicated, and one of the multiple identical transformation structures can be retained as a candidate transformation structure.

[0225] Optionally, when determining or obtaining the candidate transformation structure list according to at least one of methods a22 to a25, if the number of candidate transformation structures in the candidate transformation structure list is less than a preset number threshold (for example, the preset number threshold is 3), the number of candidate transformation structures in the candidate transformation structure list can be supplemented according to method a21 so that the number of candidate transformation structures is greater than or equal to the preset number threshold (for example, the preset number threshold is 3).

[0226] Method a3, the prediction mode for the current block;

[0227] Optionally, the prediction mode includes at least one of inter-frame prediction mode and intra-frame prediction mode.

[0228] Optionally, based on the prediction mode of the current block, it can be determined or obtained that the current block can undergo at least one transformation mode in the prediction mode. For example, in the inter-frame prediction mode, the transformation mode of the current block can include at least one of the following: DCT-2, TS, SBT, LFNST, NSPT and MTS.

[0229] Optionally, the corresponding transformation structure can be determined or obtained based on the transformation methods that the current block can perform in inter-frame prediction mode and / or intra-frame prediction mode.

[0230] Optionally, when the prediction mode of the current block is the inter-frame prediction mode, the transformation structure containing the transformation mode of non-inter-frame prediction mode in the transformation structure determined or obtained in mode a1 can be filtered out to obtain at least one transformation structure adapted to the inter-frame prediction mode.

[0231] Optionally, when the prediction mode of the current block is the intra-prediction mode, the transformation structure containing the transformation mode of non-intra-prediction mode in the transformation structure determined or obtained in mode a1 can be filtered out to obtain at least one transformation structure adapted to the intra-prediction mode.

[0232] In this approach, by specifying the transformation methods that the current block can adopt based on the prediction mode of the current block, it is possible to improve the adaptability of the transformation structure to the current block, thereby enabling improvements in the encoding and / or decoding quality during video encoding and / or decoding.

[0233] Method a4: The block size of the current block;

[0234] Optionally, the block size of the current block may include at least one of the current block's width, height, aspect ratio, perimeter, and area, and may also include variations (such as scaling down or scaling up) or indexes corresponding to at least one of the current block's width, height, aspect ratio, perimeter, and area.

[0235] Optionally, when the block size of the current block is smaller than the preset block size corresponding to the current block (hereinafter referred to as the first block size, optionally, the first size is 8), the SBT process can be skipped in the transformation mode of the current block so that the SBT contained in the determined or obtained transformation structure is a skipped SBT.

[0236] Optionally, when the size of the current block is larger than the size of the first block, the transformation method of the current block may include SBT.

[0237] Optionally, the corner mode of SBT can be determined and / or obtained based on the block size of the current block. For example, when the block size of the current block is smaller than another preset block size corresponding to the current block (hereinafter referred to as the second block size, optionally, the second size is 8), the corner mode of SBT is determined and / or obtained as not enabling the corner mode.

[0238] Optionally, when the current block size is greater than or equal to the second block size, the corner mode of the SBT is determined and / or obtained as enabled corner mode.

[0239] Optionally, the size of the second piece is greater than or equal to the size of the first piece.

[0240] Optionally, the type of corner mode enabled by SBT can be determined and / or obtained based on the current block size.

[0241] Optionally, the corner pattern types include at least a 1 / 2-width corner pattern and / or a 1 / 4-width corner pattern.

[0242] Optionally, when the current block size is smaller than another preset block size corresponding to the current block (hereinafter referred to as the third block size, optionally the third size is 16), the type of corner mode enabled by SBT is determined and / or obtained as 1 / 2 width corner mode.

[0243] Optionally, when the current block size is greater than or equal to the third block size, determine and / or obtain that the corner mode enabled by SBT is 1 / 4 width corner mode.

[0244] In this approach, by determining or obtaining the specific transformation method in at least one transformation structure based on the block size of the current block, it is possible to improve the adaptability of the transformation structure to the current block, thereby supporting the improvement of encoding and / or decoding quality in the video encoding and / or decoding process.

[0245] Method a5, the block size of at least one sub-block;

[0246] Optionally, the size of a sub-block may include at least one of the sub-block's width, height, aspect ratio, perimeter, and area, and may also include variations (such as scaling down or scaling up) or indexes corresponding to at least one of the sub-block's width, height, aspect ratio, perimeter, and area.

[0247] Optionally, when the block size of at least one sub-block of the current block is less than or equal to the preset block size corresponding to that sub-block (optionally, the preset block size is 8), the SBT process can be skipped in the transformation mode of the current block so that the SBT contained in the determined or obtained transformation structure is a skipped SBT.

[0248] Optionally, if the sum of the block sizes of the sub-blocks of the current block, and / or the variation of the sum, is greater than or equal to the preset block size of the current block, the SBT process can be skipped in the transformation mode of the current block, so that the SBT contained in the determined or obtained transformation structure is a skipped SBT.

[0249] In this approach, by determining or obtaining the specific transformation method in at least one transformation structure based on the block size of at least one sub-block, it is possible to improve the adaptability of the transformation structure to the current block, thereby supporting the improvement of encoding and / or decoding quality in the video encoding and / or decoding process.

[0250] Method a6, bitstream.

[0251] Optionally, when the processing device is an encoder, the encoder can determine or obtain the corresponding transformation structure based on the transformation configuration instructions transmitted in the code stream.

[0252] Optionally, the transformation configuration instruction is an instruction determined or generated based on the user's pre-defined or real-time transformation configuration operation, used to instruct the encoder to perform at least one transformation mode and / or at least one transformation structure during the prediction process.

[0253] Optionally, when the processing device is a decoder, the decoder can determine the corresponding transformation structure based on the transformation signaling transmitted in the bitstream.

[0254] Optionally, the transform signaling may be determined or generated by the encoder based on the transform structure used to transform the current block, and is used to indicate the transform structure used by the decoder at the encoder so that the decoder uses the same and / or opposite transform structure as the encoder to transform the transform coefficients in the bitstream.

[0255] In this approach, the transformation structure required to transform the current block can be accurately determined based on the data transmitted in the bitstream (including transformation configuration instructions and / or transformation signaling), which can support improved encoding and decoding quality during video encoding and / or decoding processes.

[0256] In this embodiment, by employing at least one of the above methods a1 to a6 to determine or obtain the transform structure, it is possible to support the use of the correlation between transform structures between spatially adjacent and / or temporally co-located regions, enabling the current block and / or sub-block to reuse the already encoded / decoded CU and / or TU transform structures. This avoids repeated signaling descriptions of the same transform structure when there is strong correlation, reduces the number of independent syntax elements in the bitstream used to indicate the transform mode, thereby reducing bit overhead and improving bitstream compression efficiency.

[0257] Third Embodiment

[0258] Based on any of the above embodiments, a third embodiment is proposed.

[0259] In this embodiment, the processing method further includes at least one of the following methods b1 to b3:

[0260] Method b1, at least one transformation structure is determined or obtained based on at least one candidate transformation structure in the candidate transformation structure list;

[0261] Optionally, at least one candidate transformation structure can be selected from the candidate transformation structure list as the transformation structure for transforming the current block, so as to transform the current block according to the selected transformation structure.

[0262] Optionally, if only one candidate transformation structure exists in the candidate transformation structure list, then that candidate transformation structure is selected as the transformation structure for transforming the current block.

[0263] Optionally, when multiple candidate transform structures exist in the candidate transform structure list, the encoder selects a candidate transform structure from the candidate transform structure list in a manner that includes at least one of the following methods b11 to b1,4:

[0264] Method b11, random selection method;

[0265] Optionally, the encoder may randomly select a candidate transform structure from the candidate transform structure list as the transform structure for transforming the current block, and transform the current block according to the randomly selected candidate transform structure.

[0266] In this approach, since the candidate transform structure list supports recording candidate transform structures that are related to the current block in the spatial dimension, temporal dimension, and / or component dimension, any candidate transform structure in the candidate transform structure list can be adapted to the current block. Therefore, any candidate transform structure can be randomly selected from the candidate transform structure list as the transform structure to transform the current block, reducing the computational complexity in the encoding and decoding process.

[0267] Method b12 selects the default transformation structure from the candidate transformation structure list;

[0268] Optionally, the encoder may select a default transform structure from the list of candidate transform structures as the transform structure for transforming the current block, so as to transform the current block according to the default transform structure.

[0269] Optionally, the order of the transformation modes in the default transformation structure can be as follows: SBT off, DCT-2, TS, SBT on + IFNST, NSPT, DCT-2 + IFNST.

[0270] In this approach, since the candidate transform structure list can contain multiple candidate transform structures, traversing each candidate transform structure involves computational complexity. Therefore, a default transform structure can be selected from the candidate transform structure list as the transform structure for transforming the current block, thereby reducing the computational complexity during the encoding and decoding process.

[0271] Method b13 selects the candidate transform structure with the lowest rate-distortion cost from the candidate transform structure list;

[0272] Optionally, the encoder can perform rate-distortion optimization on each candidate transform structure recorded in the candidate transform structure list, determine or obtain the rate-distortion cost of each candidate transform structure, and select the candidate transform structure with the smallest rate-distortion cost among the candidate transform structures as the transform structure for transforming the current block, so as to transform the current block according to the candidate transform structure with the smallest rate-distortion cost.

[0273] Optionally, when the minimum rate distortion cost corresponds to at least one different candidate transformation structure, any one of the at least one different candidate transformation structures can be selected as the transformation structure for transforming the current block, so as to transform the current block according to the candidate transformation structure with the minimum rate distortion cost.

[0274] In this approach, by optimizing each candidate transform structure in the candidate transform structure list through traversal rate distortion, the candidate transform structure with the highest fit to the current block can be determined or obtained. This avoids performing traversal rate distortion optimization on the full transform combination that fits the current block, thereby reducing computational complexity and / or supporting improvements in the encoding and / or decoding quality during video encoding and / or decoding.

[0275] Method b14 selects a candidate transformation structure at a specified position in the candidate transformation structure list.

[0276] Optionally, the specified position can be the list index of the candidate transformation structure list and / or the order of the candidate transformation structures in the candidate transformation structure list.

[0277] Optionally, the user can pre-set and / or set a specified position to select the corresponding candidate transformation structure from the candidate transformation structure list based on the specified position. For example, if the specified position set by the user is the Nth candidate transformation structure recorded in the candidate transformation structure list (N is any integer from 1 to M, M is the number of candidate transformation structures recorded in the candidate transformation structure list, and M is greater than 1), then the Nth candidate transformation structure will be used as the transformation structure to transform the current block, so that the current block can be transformed based on the Nth candidate transformation structure.

[0278] Optionally, since at least one candidate transformation structure recorded in the candidate transformation structure list can be sorted within the candidate transformation structure list according to its fit with the current block, a candidate transformation structure can be selected according to the sorting method of the candidate transformation structure list and a specified position. For example, when the sorting method is to sort according to the fit from high to low, the specified position can be the first candidate transformation structure recorded in the candidate transformation structure list, so that the first candidate transformation structure in the candidate transformation structure list that has the highest fit with the current block is used as the transformation structure to transform the current block.

[0279] Optionally, when the sorting method is to sort according to the fit from low to high, the specified position can be the last candidate transformation structure recorded in the candidate transformation structure list, so that the last candidate transformation structure in the candidate transformation structure list with the highest fit with the current block is used as the transformation structure to transform the current block.

[0280] In this approach, by selecting a candidate transform structure at a specified position in the candidate transform structure list, it is possible to select the candidate transform structure list with the highest fit to the current block. This avoids the complex rate-distortion cost calculation process for the candidate transform structures in the candidate transform structure list, and / or avoids selecting candidate transform structures with limited fit to the current block from the candidate structure list based on random selection. This reduces computational complexity and / or supports improved encoding and / or decoding quality in the video encoding and / or decoding process.

[0281] Method b2, at least one transformation structure is determined or obtained based on the transformation signaling in the code stream;

[0282] Optionally, the decoder can determine or obtain at least one transform structure based on the transform signaling in the bitstream.

[0283] Optionally, since the decoder can construct the same list of candidate transform structures as the encoder according to the same construction scheme, the decoder can determine or obtain the same and / or opposite transform structures as the encoder from the same list of candidate transform structures based on the transform signaling generated by the encoder, so as to transform the transform coefficients in the bitstream.

[0284] Optionally, the decoder can determine the candidate transform structure corresponding to the list index in the transform signaling as the transform structure to transform the current block.

[0285] In this method, the decoder can determine or obtain at least one transformation structure that is the same as or relative to the encoder based on the transformation signaling in the bitstream. This allows the decoder to accurately determine the transformation structure required to transform the current block, thereby supporting the improvement of encoding and / or decoding quality in the video encoding and / or decoding process.

[0286] Method b3, the list index of at least one candidate transformation structure is determined or obtained based on the sorting method of the candidate transformation structure list and / or the sorting position of at least one candidate transformation structure.

[0287] Optionally, the sorting method is a predefined rule or algorithm used to organize and arrange the candidate transformation structures in the candidate transformation structure list.

[0288] Optionally, the sorting method may be based on the chronological order in which the candidate transform structures are recorded in the candidate transform structure list, and / or based on the block size similarity between the current block and candidate blocks (which may include at least one of the following: adjacent block above, non-adjacent block above, adjacent block to the left, non-adjacent block to the left, adjacent block to the upper left, non-adjacent block to the upper left, cross-component block, same-component block, co-located block, and time-domain block), and / or based on the sub-blocks of the current block and candidate sub-blocks (which may include adjacent sub-blocks above, non-adjacent sub-blocks above, adjacent sub-blocks to the left, non-adjacent sub-blocks to the left, adjacent sub-blocks to the upper left, non-adjacent sub-blocks to the upper left, cross-component block, and cross-component block). The similarity of sub-block sizes (at least one of sub-blocks, co-occurring sub-blocks, co-occurring sub-blocks, and time-domain sub-blocks) and / or the reuse frequency of multiplexed transform structures are sorted in ascending or descending order. The purpose is to make candidate transform structures that are more likely to be selected by the current block appear at the top of the candidate transform structure list, so that the list index of the candidate transform structure selected in the candidate transform structure list is a smaller value. This allows the encoder to support a reduction in the number of encoded bits during the encoding of the list index, thereby reducing the bit overhead related to the candidate transform structure in the bitstream and thus supporting an improvement in bitstream compression efficiency.

[0289] Optionally, the encoder may encode the list index using at least one of unary code encoding, truncated unary code encoding, and exponential Golomb encoding.

[0290] Optionally, when the encoding method is unary code encoding, the unary code encoding results of list indices "0" to "5" are "0", "10", "110", "1110", "11110" and "111110" respectively. For example, when the candidate transform structure selected in the candidate transform structure list is the second candidate transform structure, the unary code encoding result of the second candidate transform structure corresponding to list index "2" is "110".

[0291] Optionally, when the encoding method is truncated unary code encoding, the truncated unary code encoding results of list indices "0" to "5" are "0", "10", "110", "1110", "11110" and "111111" respectively. For example, when the candidate transform structure selected in the candidate transform structure list is the second candidate transform structure, the truncated unary code encoding result of the second candidate transform structure corresponding to list index "2" is "110".

[0292] Optionally, when the encoding method is exponential Golomb encoding, the exponential Golomb encoding results of list indices "0" to "5" are "0", "100", "101", "11000", "11001" and "11010" respectively. For example, when the candidate transform structure selected in the candidate transform structure list is the second candidate transform structure, the exponential Golomb encoding result of the second candidate transform structure corresponding to list index "2" is "101".

[0293] Optionally, the sorting position is the order in which each candidate transformation structure is located in the candidate transformation structure list formed according to the sorting method.

[0294] Optionally, the sorting position of the candidate transformation structure in the first order in the candidate transformation structure list can be 1, the sorting position of the candidate transformation structure in the second order can be 2, the sorting position of the candidate transformation structure in the third order can be 3, and so on.

[0295] Optionally, the sorting position can be used to indicate the value of the list index of the candidate transformation structure.

[0296] Optionally, the list index may include at least one of the value of the sort position and the mapping value of the sort position.

[0297] Optionally, the list index of at least one candidate transformation structure can be determined or obtained based on the sorting method of the candidate transformation structure list and the sorting position of at least one candidate transformation structure.

[0298] Optionally, when the sorting method is to sort the candidate transformation structures in descending order based on the block size similarity between the current block and the candidate blocks, the sorting position of the candidate transformation structure in the first order in the candidate transformation structure list can be 1, the sorting position of the candidate transformation structure in the second order can be 2, the sorting position of the candidate transformation structure in the third order can be 3, and so on. The list index of each candidate transformation structure determined or obtained in this way is the value of the sorting position minus 1.

[0299] Optionally, when the sorting method is to sort the candidate transformation structures in ascending order based on the similarity of the block size between the current block and the candidate blocks, the sorting position of the candidate transformation structure in the first order in the candidate transformation structure list can be 1, the sorting position of the candidate transformation structure in the second order can be 2, the sorting position of the candidate transformation structure in the third order can be 3, and so on. The list index of each candidate transformation structure determined or obtained in this way is the mapping value of the sorting position value. For example, the mapping relationship can be 1 mapped to N-1 (N is the number of candidate transformation structures in the candidate transformation structure list, N is greater than 1), 2 mapped to N-2... N-2 mapped to 2, and N-1 mapped to 1.

[0300] Optionally, when the sorting method is to sort the candidate transformation structures in descending order based on the reuse frequency of the current block and the multiplexing transformation structure, the sorting position of the candidate transformation structure in the first order in the candidate transformation structure list can be 1, the sorting position of the candidate transformation structure in the second order can be 2, the sorting position of the candidate transformation structure in the third order can be 3, and so on. The list index of each candidate transformation structure determined or obtained in this way is the value of the sorting position minus 1.

[0301] Optionally, when the sorting method is to sort the candidate transformation structures in ascending order based on the reuse frequency of the current block and the multiplexing transformation structure, the sorting position of the candidate transformation structure in the first order in the candidate transformation structure list can be 1, the sorting position of the candidate transformation structure in the second order can be 2, the sorting position of the candidate transformation structure in the third order can be 3, and so on. The list index of each candidate transformation structure determined or obtained in this way is the mapping value of the sorting position value. For example, the mapping relationship can be 1 mapped to N-1 (N is the number of candidate transformation structures in the candidate transformation structure list, N is greater than 1), 2 mapped to N-2...N-2 mapped to 2, and N-1 mapped to 1.

[0302] In this approach, by determining or obtaining the list index based on the sorting method of the candidate transform structure list and / or the sorting position of at least one candidate transform structure, the number of encoded bits can be reduced during the encoder's encoding of the list index, thereby reducing the bit overhead related to the candidate transform structure in the bitstream and thus improving the bitstream compression efficiency.

[0303] Optionally, the sorting position of at least one candidate transformation structure is determined or obtained according to at least one of the following methods b31 to b36:

[0304] Method b31, the block size of the current block;

[0305] Optionally, the block size of the current block may include at least one of the current block's width, height, aspect ratio, perimeter, and area, and may also include variations (such as scaling down or scaling up) or indexes corresponding to at least one of the current block's width, height, aspect ratio, perimeter, and area.

[0306] Optionally, the block size similarity between the current block and each candidate block is determined or obtained based on the block size of the current block, and the candidate transformation structures corresponding to each candidate block are sorted according to the block size similarity and sorting method, so as to determine or obtain the sorting position of each candidate transformation structure in the candidate transformation structure list.

[0307] Alternatively, the block size similarity (Dsize1) can be represented as at least one of the following:

[0308] Dsize1 = |w_cand1 - w_cur1| + |h_cand1 - h_cur1|;

[0309] Dsize1 = max(|w_cand1 - w_cur1|, |h_cand1 - h_cur1|);

[0310] w_cand1 is the width of the candidate block, w_cur1 is the width of the current block, h_cand1 is the height of the candidate block, and h_cur1 is the height of the current block. Block size similarity can reflect the size difference between the current block and the candidate block.

[0311] In this approach, by determining or obtaining the sorting position of each candidate transform structure in the candidate transform structure list based on the block size of the current block, the candidate transform structures corresponding to candidate blocks that are geometrically closer to the current block can be selected first, thereby increasing the hit probability of candidate transform structure reuse and thus supporting improved encoding and / or decoding efficiency in the video encoding and / or decoding process.

[0312] Method b32, the block size of at least one sub-block;

[0313] Optionally, the block size of at least one sub-block may include at least one of the following: width, height, aspect ratio, perimeter, and area of ​​the sub-block. It may also include variations (such as scaling down or scaling up) or indexes corresponding to at least one of the following: width, height, aspect ratio, perimeter, and area of ​​the sub-block.

[0314] Optionally, the similarity of the size of at least one sub-block with each candidate sub-block is determined or obtained based on the block size of at least one sub-block, and the candidate transformation structures corresponding to each candidate sub-block are sorted according to the similarity of the size of each sub-block and the sorting method, so as to determine or obtain the sorting position of each candidate transformation structure in the candidate transformation structure list.

[0315] Alternatively, the sub-block size similarity (Dsize2) can be expressed as at least one of the following:

[0316] Dsize2= |w_cand2 - w_cur2| + |h_cand2 - h_cur2|;

[0317] Dsize2 = max(|w_cand2 - w_cur2|, |h_cand2 - h_cur2|);

[0318] w_cand2 is the width of the candidate sub-block, w_cur2 is the width of the sub-block, h_cand2 is the height of the candidate sub-block, and h_cur2 is the height of the sub-block. The sub-block size similarity can reflect the size difference between the sub-block and the candidate sub-block.

[0319] In this approach, by determining or obtaining the sorting position of each candidate transform structure in the candidate transform structure list based on the block size of at least one sub-block, the candidate transform structures corresponding to candidate sub-blocks that are geometrically closer to the sub-block can be selected first, thereby increasing the hit probability of candidate transform structure reuse and thus supporting improved encoding and / or decoding efficiency in the video encoding and / or decoding process.

[0320] Method b33, the block size of at least one of the following: the upper adjacent block, the upper non-adjacent block, the left adjacent block, the left non-adjacent block, the upper left adjacent block, the upper left non-adjacent block, the cross-component block, the same component block, the same position block, and the time domain block;

[0321] Optionally, the block size of at least one of the following: the upper adjacent block, the upper non-adjacent block, the left adjacent block, the left non-adjacent block, the upper left adjacent block, the upper left non-adjacent block, the cross-component block, the same component block, the same position block, and the time domain block (hereinafter referred to as the candidate block) may include at least one of the following: the width, height, aspect ratio, perimeter, and area of ​​the candidate block. It may also include a variation (such as shrinking or enlarging) or index corresponding to at least one of the following: the width, height, aspect ratio, perimeter, and area of ​​the candidate block.

[0322] Optionally, the block size similarity between the candidate block and the current block is determined or obtained based on the block size of the candidate block, and the candidate transformation structures corresponding to each candidate block are sorted according to the block size similarity and sorting method of each candidate block, so as to determine or obtain the sorting position of each candidate transformation structure in the candidate transformation structure list.

[0323] Alternatively, the block size similarity (Dsize1) can be represented as at least one of the following:

[0324] Dsize1 = |w_cand1 - w_cur1| + |h_cand1 - h_cur1|;

[0325] Dsize1 = max(|w_cand1 - w_cur1|, |h_cand1 - h_cur1|);

[0326] w_cand1 is the width of the candidate block, w_cur1 is the width of the current block, h_cand1 is the height of the candidate block, and h_cur1 is the height of the current block. Block size similarity can reflect the size difference between the current block and the candidate block.

[0327] In this approach, by determining or obtaining the sorting position of each candidate transform structure in the candidate transform structure list based on the block size of the candidate block, the candidate transform structure corresponding to the candidate block that is geometrically closer to the current block can be selected first, thereby increasing the hit probability of candidate transform structure reuse and thus supporting the improvement of encoding and / or decoding efficiency in the video encoding and / or decoding process.

[0328] Method b34, the block size of at least one of the following: the upper adjacent sub-block, the upper non-adjacent sub-block, the left adjacent sub-block, the left non-adjacent sub-block, the upper left adjacent sub-block, the upper left non-adjacent sub-block, the cross-component sub-block, the same component sub-block, the same position sub-block, and the time domain sub-block;

[0329] Optionally, the block size of at least one of the following sub-blocks (hereinafter referred to as candidate sub-blocks): the upper adjacent sub-block, the upper non-adjacent sub-block, the left adjacent sub-block, the left non-adjacent sub-block, the upper left adjacent sub-block, the upper left non-adjacent sub-block, the cross-component sub-block, the same component sub-block, the co-position sub-block, and the time domain sub-block (hereinafter referred to as candidate sub-blocks) may include at least one of the following: the width, height, aspect ratio, perimeter, and area of ​​the candidate sub-block. It may also include a variation (such as shrinking or enlarging) or index corresponding to at least one of the following: the width, height, aspect ratio, perimeter, and area of ​​the candidate sub-block.

[0330] Optionally, the similarity between the sub-block size of at least one candidate sub-block and the sub-block of the current block is determined or obtained based on the block size of at least one candidate sub-block, and the candidate transformation structures corresponding to each candidate sub-block are sorted according to the sub-block size similarity and sorting method of each candidate sub-block, so as to determine or obtain the sorting position of each candidate transformation structure in the candidate transformation structure list.

[0331] Alternatively, the sub-block size similarity (Dsize2) can be expressed as at least one of the following:

[0332] Dsize 2= |w_cand2 - w_cur2| + |h_cand2 - h_cur2|;

[0333] Dsize2 = max(|w_cand2 - w_cur2|, |h_cand2 - h_cur2|);

[0334] w_cand2 is the width of the candidate sub-block, w_cur2 is the width of the sub-block, h_cand2 is the height of the candidate sub-block, and h_cur2 is the height of the sub-block. The sub-block size similarity can reflect the size difference between the sub-block and the candidate sub-block.

[0335] In this approach, by determining or obtaining the sorting position of each candidate transform structure in the candidate transform structure list based on the block size of at least one candidate sub-block, the candidate transform structures corresponding to candidate sub-blocks that are geometrically closer to the sub-block can be selected first, thereby increasing the hit probability of candidate transform structure reuse and thus supporting improved encoding and / or decoding efficiency in the video encoding and / or decoding process.

[0336] Method b35, the reuse frequency of at least one candidate transform structure.

[0337] Optionally, the reuse frequency is a statistical information on the number of times at least one candidate transform structure appears repeatedly in the candidate transform structure list during the construction of the candidate transform structure list, and / or in the encoded / decoded candidate blocks associated with the current block, and / or in the encoded / decoded candidate sub-blocks associated with at least one sub-block of the current block.

[0338] Optionally, in the candidate transformation structure list, if the same candidate transformation structure is used by multiple candidate blocks from different spatial or temporal dimensions, then the reuse frequency of the candidate transformation structure is the number of candidate blocks that use the candidate transformation structure.

[0339] Optionally, the reuse frequency can be used to determine and / or adjust the sorting position of the candidate transform structure in the candidate transform structure list.

[0340] Optionally, the candidate transformation structures corresponding to the candidate sub-blocks are sorted according to the reuse frequency of at least one candidate sub-block, so as to determine or obtain the sorting position of each candidate transformation structure in the candidate transformation structure list.

[0341] Optionally, each candidate transformation structure in the candidate transformation structure list is mapped to a corresponding structural combination key, and the reuse frequency of each candidate transformation structure is determined or obtained based on the statistical frequency of each structural combination key.

[0342] Alternatively, a structural key can be represented as:

[0343] StructKey=(sbt_info,mts_idx,idxLFNST);

[0344] sbt_info is used to characterize the SBT's enabled status and SBT's segmentation method information. mts_idx is the index of the main transform kernel (including DCT-2, TS and / or MTS) of the candidate transform structure corresponding to the structural combination key. idxLFNST is the index of the non-separable transform (including LFNST and / or NSPT).

[0345] In this approach, by determining or obtaining the sorting position of each candidate transform structure in the candidate transform structure list based on the reuse frequency of at least one candidate transform structure, the candidate transform structure with higher usage frequency in the prediction process of the current block can be selected first, increasing the hit probability of candidate transform structure reuse. And / or the statistical process and reordering process of the reuse frequency only rely on the information already in the candidate transform structure list. Both the encoder and / or decoder can complete the same statistical process and reordering process without introducing additional syntax elements, achieving reproducibility at the decoder end, and thus supporting the improvement of encoding and / or decoding efficiency in the video encoding and / or decoding process.

[0346] Optionally, when at least one of methods b31 to b35 in this embodiment is used to reorder the sorting position of each candidate transformation structure in the candidate transformation structure list, the execution order of at least one of methods b31 to b34 can be before method b35 (i.e., execution order S1: first reorder based on size similarity, and then reorder a second time based on reuse frequency on the candidate transformation structure list obtained by reordering based on size similarity), or after method b35 (i.e., execution order S2: first reorder based on reuse frequency, and then reorder a second time based on size similarity on the candidate transformation structure list obtained by reordering based on reuse frequency).

[0347] Optionally, during the secondary reordering process, if execution order S1 is adopted, then based on the candidate transformation structure list obtained by reordering based on size similarity, the sorting positions of at least one candidate transformation structure with the same size similarity in the candidate transformation structure list are reordered based on reuse frequency.

[0348] Optionally, during the secondary reordering process, if execution order S2 is adopted, then based on the candidate transformation structure list obtained by reordering based on reuse frequency, the sorting positions of at least one candidate transformation structure with the same reuse frequency in the candidate transformation structure list are reordered based on size similarity.

[0349] Optionally, during the secondary reordering process, if at least one candidate transformation structure with the same size similarity has the same reuse frequency, the candidate transformation structure list obtained by reordering based on size similarity can be used as the candidate transformation structure list obtained by secondary reordering.

[0350] Optionally, during the secondary reordering process, if at least one candidate transform structure with the same reuse frequency has the same size similarity, the candidate transform structure list obtained by reordering based on reuse frequency can be used as the candidate transform structure list obtained by secondary reordering.

[0351] Optionally, during and / or after the secondary reordering, at least one identical candidate transformation structure in the candidate transformation structure list can be deduplicated, and one candidate transformation structure can be retained among the at least one identical candidate transformation structure.

[0352] In this embodiment, the candidate transform structure list is reordered based on the block size similarity between the current block and the candidate blocks, and / or the reuse frequency of the candidate transform structure in the list. This allows the candidate transform structure that is statistically more likely to be selected by the current block to obtain a higher sorting position (i.e., a smaller list index value). When the list index is encoded using, for example, a context-adaptive truncated unary code encoding method, a smaller index value corresponds to a shorter binary codeword, thereby reducing the average number of encoded bits for the transform signaling, and thus reducing the bit overhead of the bit stream and improving the bit stream compression efficiency.

[0353] Fourth embodiment

[0354] Based on any of the above embodiments, a fourth embodiment is proposed.

[0355] In this embodiment, the transformation signaling is determined or obtained according to at least one of the following methods c1 to c5:

[0356] Method c1, transform structure identifier;

[0357] Optionally, the transform structure identifier can be in the form of a transform structure used to indicate in the bitstream that the encoder uses a transform scheme for the current block and / or a conventional independent transform scheme.

[0358] Optionally, the transformation structure identifier can be a flag (e.g., tr_merge_flag).

[0359] Optionally, when the flag is set to the preset first character (e.g., tr_merge_flag=1), it indicates that the encoder uses only the transformation structure form for the current block.

[0360] Optionally, when the flag bit is a preset second character (e.g., tr_merge_flag=0), it indicates that the encoder uses only the traditional independent transformation method for the current block.

[0361] Optionally, when the flag is a preset third character (e.g., tr_merge_flag=11), it indicates that the encoder uses a transformation scheme for the current block, including the form of a transformation structure and the form of a traditional independent transformation.

[0362] Optionally, when the flag is a preset third character, the encoder does not include the traditional independent transformation method in the transformation structure used for the current block.

[0363] Optionally, the transformation structure identifier can be used as a component of the transformation signaling, and the transformation signaling can be constructed together with the transformation structure identifier and the other components.

[0364] Optionally, the remaining components may include at least one of the list index encoding result, the syntax element of the transformation mode, and the transformation kernel encoding result.

[0365] In this approach, by constructing transform signaling based on transform structure identifiers, the encoder can accurately determine the transform form for transforming the current block in the encoder, thereby supporting improvements in the encoding and / or decoding quality during video encoding and / or decoding.

[0366] Optionally, at least one transformation structure identifier is determined or obtained according to at least one of the following methods c11 and c12:

[0367] Method c11, the number of candidate transformation structures;

[0368] Optionally, at least one transformation structure identifier is determined or obtained based on the number of candidate transformation structures;

[0369] Optionally, when the number of candidate transform structures is a preset first number (e.g., 0), indicating that the candidate transform structure list does not exist and / or no candidate transform structure is recorded in the candidate transform structure list, a preset second character (e.g., 0) can be used as the transform structure identifier to indicate that the encoder transforms the current block only according to the traditional independent transform method, and to instruct the decoder to transform the current block only according to the traditional independent transform method (i.e., inverse transform).

[0370] Optionally, when the number of candidate transform structures is a preset second number (e.g., 1), it indicates that there is only one candidate transform structure in the candidate transform structure list. Since there is also only one unique and identical candidate transform structure in the decoder's candidate transform structure list, the list index encoding result can be omitted when constructing the transform signaling. A preset first character (e.g., 1) or a preset third character (e.g., 11) can be used as the transform structure identifier.

[0371] Optionally, if the number of candidate transformation structures is greater than a preset second number (e.g., 1), it indicates that there are multiple candidate transformation structures in the candidate transformation structure list, and a preset first character (e.g., 1) or a preset third character (e.g., 11) can be used as the transformation structure identifier.

[0372] Method c12, transformation cost.

[0373] Optionally, the transformation cost may be a first rate-distortion cost for transforming the current block according to at least one transformation structure, and / or a second rate-distortion cost for transforming the current block according to at least one conventional independent transformation scheme.

[0374] Optionally, the encoder can determine or obtain at least one transform structure identifier based on the first rate distortion cost and the second rate distortion cost.

[0375] Optionally, when the first rate-distortion cost is less than or equal to the second rate-distortion cost, indicating that the transformation cost of transforming the current block according to at least one transformation structure is less than or equal to the transformation cost of transforming the current block according to at least one conventional independent transformation method, a preset first character (e.g., 1) or a preset third character (e.g., 11) can be used as the transformation structure identifier to indicate that the encoder transforms the current block according to at least one transformation structure, and to instruct the decoder to transform the current block according to at least one transformation structure (i.e., inverse transformation).

[0376] Optionally, when the first rate distortion cost is greater than the second rate distortion cost, indicating that the transformation cost of transforming the current block according to at least one transform structure is greater than the transformation cost of transforming the current block according to at least one conventional independent transform method, a preset second character (e.g., 0) can be used as the transform structure identifier to indicate that the encoder transforms the current block only according to the conventional independent transform method, and to instruct the decoder to transform the current block only according to the conventional independent transform method (i.e., inverse transform).

[0377] Optionally, in the process of determining or obtaining at least one transformation structure identifier, the necessity of transforming the current block according to the transformation structure can be determined according to this method (i.e., it is necessary when the first rate distortion cost is less than or equal to the second rate distortion cost, and it is not necessary when the first rate distortion cost is greater than the second rate distortion cost).

[0378] Optionally, when necessary, at least one transformation structure identifier may be determined or obtained according to method c11.

[0379] Optionally, when it is not necessary, it is possible to skip determining or obtaining at least one transformation structure identifier according to method c11 and use the preset second character (e.g., 0) determined or obtained by this method as the transformation structure identifier.

[0380] Method c2, list index encoding result;

[0381] Optionally, the list index encoding result can be the encoding result obtained by the encoder encoding the list index using a preset encoding method (e.g., tr_merge_idx).

[0382] Optionally, the preset encoding method can be CABAC encoding, which may include at least one of unary code encoding, truncated unary code encoding, and exponential Golomb encoding.

[0383] Optionally, at least one list index encoding result includes at least one of the following: a first unary code encoding result, a first truncated unary code encoding result, and a first exponential Columbus encoding result;

[0384] Optionally, the first unary code encoding structure can be the encoding result obtained by the encoder performing unary code encoding on the list index. For example, the first unary code encoding results of list indices "0" to "5" are "0", "10", "110", "1110", "11110" and "111110" respectively.

[0385] Optionally, the first truncated unary code encoding result can be the encoding result obtained by the encoder by truncating the list index into unary code. For example, the truncated unary code encoding results of list indices "0" to "5" are "0", "10", "110", "1110", "11110" and "111111" respectively.

[0386] Optionally, the first exponential Golomb encoding result can be the encoding result obtained by the encoder performing exponential Golomb encoding on the list index. For example, the exponential Golomb encoding results of list indices "0" to "5" are "0", "100", "101", "11000", "11001" and "11010" respectively.

[0387] Optionally, the list index encoding result is used as a component of the transformation signaling, and the transformation signaling is constructed together with the list index encoding result and the other components.

[0388] Optionally, the remaining components may include at least one of mode c1, the syntax elements of the transformation mode, and the transformation kernel encoding result.

[0389] Optionally, when the transformation signaling is determined or obtained solely based on the transformation structure identifier and the list index encoding result, the transformation structure identifier may be a preset first character (e.g., 1).

[0390] Optionally, when determining or obtaining the transformation signaling based on at least one of the syntax elements of the transformation mode and the encoding result of the transformation kernel, and the encoding result of the transformation structure identifier and the list index, the transformation structure identifier can be a preset third character (e.g., 11).

[0391] In this approach, by constructing transform signaling based on the list index encoding results, the encoder can accurately determine the transform structure that transforms the current block in the encoder. The transform form is not described, transmitted, and parsed independently by the traditional independent transform method, but is determined all at once by the list index encoding results. This can reduce the bit overhead of the transform method that transforms the current block in the bitstream, thereby supporting improved bitstream compression efficiency.

[0392] Method c3, the number of candidate transformation structures;

[0393] Optionally, at least one transformation signaling may be determined or obtained based on the number of candidate transformation structures.

[0394] Optionally, when the number of candidate transform structures is a preset first number (e.g., 0), indicating that the candidate transform structure list does not exist and / or no candidate transform structure is recorded in the candidate transform structure list, it can be determined that the transform signaling may not include the list index encoding result, so that the transform signaling can be determined or obtained only based on at least one of the transform structure identifier, the syntax element of the transform mode, and the transform kernel encoding result.

[0395] Optionally, when the number of candidate transform structures is a preset second number (e.g., 1), it indicates that there is only one unique candidate transform structure in the candidate transform structure list. Since there is also only one unique and identical candidate transform structure in the decoder's candidate transform structure list, it can be determined that the transform signaling may not contain the list index encoding result, so that the transform signaling can be determined or obtained based on at least one of the transform structure identifier, the syntax element of the transform mode, and the transform kernel encoding result.

[0396] In this approach, by constructing transform signaling based on the number of candidate transform structures, the encoder can accurately determine the transform form for transforming the current block in the encoder, thereby supporting improvements in the encoding and / or decoding quality during video encoding and / or decoding.

[0397] c4 is a syntax element for changing the mode;

[0398] Optionally, at least one transformation method includes at least one of the following: DCT-2, TS, SBT, LFNST, NSPT, and MTS.

[0399] Optionally, the syntax elements of the transformation scheme can be syntax elements of a traditional independent transformation scheme. For example, for a traditional independent SBT, its syntax elements may include at least one of the following: subblock transformation identifier (e.g., sbt_flag), 1 / 2CU width corner pattern (e.g., sbt_quad_flag), 1 / 4CU width corner pattern (e.g., sbt_quader_flag), corner selection index (e.g., horIdx / verIdx), rectangle pattern identifier (e.g., sbtQuadFlag), coding unit segmentation direction identifier (e.g., sbtHorFlag), and rectangle subblock position identifier (e.g., sbt_pos_flag).

[0400] Optionally, the syntax element for transformation scheme is used to indicate the transformation scheme adopted by the encoder for the current block, including the form of traditional independent transformation scheme.

[0401] Optionally, the syntax elements of the transformation mode are used as components of the transformation signaling, and the transformation signaling is constructed together with the syntax elements of the transformation mode and the remaining components.

[0402] Optionally, the remaining components may include at least one of mode c1, mode c2, and the transform kernel encoding result.

[0403] Optionally, when the transformation signaling is determined or obtained based solely on at least one of the syntax elements of the transformation mode and the encoding result of the transformation kernel and the transformation structure identifier, the transformation structure identifier may be a preset second character (e.g., 0).

[0404] Optionally, when determining or obtaining the transformation signaling based on at least one of the syntax elements of the transformation mode and the encoding result of the transformation kernel, and the encoding result of the transformation structure identifier and the list index, the transformation structure identifier can be a preset third character (e.g., 11).

[0405] In this approach, by determining or obtaining the transformation signaling based on the syntax elements of the transformation method, the encoder can accurately determine the transformation method for transforming the current block in the encoder, thereby supporting improvements in the encoding and / or decoding quality during video encoding and / or decoding.

[0406] Method c5 transforms the kernel encoding result.

[0407] Optionally, the transformation kernel encoding result can be the encoding result obtained by the encoder encoding the index in the syntax element of the transformation mode using a preset encoding method (e.g., mts_idx and / or idxLFNST).

[0408] Optionally, the preset encoding method can be CABAC encoding, which may include at least one of unary code encoding, truncated unary code encoding, and exponential Golomb encoding.

[0409] Optionally, at least one transform kernel encoding result includes at least one of the following: a second unary code encoding result, a second truncated unary code encoding result, and a second exponential Columbus code result.

[0410] Optionally, the second unary code encoding structure can be the encoding result obtained by the encoder performing unary code encoding on the indexes in the syntax elements of the transformation mode. It can include the unary code encoding result of the main transform kernel index (e.g., mts_idx under unary code encoding) and / or the unary code encoding result of the inseparable transform index (e.g., idxLFNST under unary code encoding).

[0411] Optionally, the second truncated unary code encoding result can be the encoding result obtained by the encoder truncating the index in the syntax element of the transform mode using unary code encoding. It can include the truncated unary code encoding result of the main transform kernel index (e.g., mts_idx under truncated unary code encoding) and / or the truncated unary code encoding result of the inseparable transform index (e.g., idxLFNST under truncated unary code encoding).

[0412] Optionally, the second exponential Golomb encoding result can be the encoding result obtained by the encoder performing exponential Golomb encoding on the index in the syntax element of the transform mode. For example, it can include the exponential Golomb encoding result of the main transform kernel index (e.g., mts_idx under exponential Golomb encoding) and / or the exponential Golomb encoding result of the non-separable transform index (e.g., idxLFNST under exponential Golomb encoding).

[0413] Optionally, the transformation kernel encoding result can be used as a component of the transformation signaling, and the transformation signaling can be constructed together with the transformation kernel encoding result and the other components.

[0414] Optionally, the remaining components may include at least one of mode c1, mode c2 and mode c4.

[0415] Reference Figure 11 During the transformation of the current block and / or its sub-blocks, if the encoder uses the candidate transformation structure proposed in this embodiment to transform the current block, then the transformation structure identifier is encoded, a candidate transformation structure list is constructed, and during and / or after the construction of the candidate transformation structure list, the candidate transformation structures in the candidate transformation structure list are checked for legality and deduplicated according to the preset conditions of the current block. The list index of the selected candidate transformation structure is determined by rate-distortion optimization, so as to apply the target transformation kernel to quantize the transformation coefficients, and the list index is encoded to obtain the list index encoding result. And / or if the encoder uses the traditional independent transformation method to transform the current block, then the transformation structure is determined according to the traditional independent transformation method, so as to apply the target transformation kernel to quantize the transformation coefficients.

[0416] Reference Figure 12 During the transformation process after the decoder parses the bitstream and performs entropy decoding, the decoder determines whether to use the candidate transform structure proposed in this embodiment to transform the current block or to use the traditional independent transform method based on the transform structure identifier. If the decoder uses the transform structure proposed in this embodiment to transform the current block, the decoder constructs a candidate transform structure list. During and / or after the construction of the candidate transform structure list, the decoder performs a validity check and deduplication on the candidate transform structures in the candidate transform structure list according to the preset conditions of the current block. The decoder decodes the list index encoding result to determine the candidate transform structure. The transform kernel of the candidate transform structure is used to perform an inverse transform on the transform coefficients obtained by inverse quantization to obtain the residual. Or, if the decoder uses the traditional independent transform method to transform the current block, the decoder determines the transform structure according to the traditional independent transform method. The transform kernel of the transform structure is used to perform an inverse transform on the transform coefficients obtained by inverse quantization to obtain the residual.

[0417] Optionally, when the transformation signaling is determined or obtained based solely on at least one of the syntax elements of the transformation mode and the encoding result of the transformation kernel and the transformation structure identifier, the transformation structure identifier may be a preset second character (e.g., 0).

[0418] Optionally, when determining or obtaining the transformation signaling based on at least one of the syntax elements of the transformation mode and the encoding result of the transformation kernel, and the encoding result of the transformation structure identifier and the list index, the transformation structure identifier can be a preset third character (e.g., 11).

[0419] In this approach, by determining or obtaining the transformation signaling based on the transformation kernel encoding result, the encoder can accurately determine the transformation method for transforming the current block in the encoder, thereby supporting the improvement of encoding and / or decoding quality in the video encoding and / or decoding process.

[0420] Optionally, to aid in understanding the implementation flow of the processing method obtained by combining the first and / or second and / or third embodiments described above, please refer to... Figure 13For the encoder, given the current block and / or at least one sub-block of the current block, determine whether to use the candidate transform structure proposed in this embodiment to transform the current block. If not, perform the traditional transform encoding process. If yes, initialize the candidate transform structure list, add adjacent sub-blocks, non-adjacent sub-blocks, and co-position sub-blocks as candidate transform structures in sequence, perform legality checks and deduplication, and reorder the candidate transform structure list based on size similarity priority. Determine whether the number of candidate transform structures is greater than or equal to a preset number threshold (e.g., the preset number threshold is 3). If yes, select the optimal candidate transform structure through rate-distortion optimization, and encode the list index of the candidate transform structure using truncated unary code to write the list index change result into the bitstream. And / or, if not, add the default transform structure to the candidate transform structure list, perform legality checks and deduplication, select the optimal candidate transform structure through rate-distortion optimization, and encode the list index of the candidate transform structure using truncated unary code to write the list index change result into the bitstream.

[0421] Alternatively, please refer to Figure 14 For the decoder, the input is an encoded bitstream (i.e., a bitstream). The decoder performs at least one of the following: decoding the bitstream, predicting relevant syntax and / or residual coefficients, and / or parsing the transform structure identifier. It determines whether the transform structure identifier is a preset first character (e.g., 0). If not, it performs the traditional transform decoding process. If yes, it initializes a candidate transform structure list, sets an upper limit and a preset threshold (e.g., a preset threshold of 3) for the number of candidate transform structures it can hold, and sequentially adds adjacent sub-blocks, non-adjacent sub-blocks, and co-position sub-blocks as candidate transform structures. It performs validity checks and deduplication, and reorders the candidate transform structure list based on size similarity. It determines whether the number of candidate transform structures is greater than or equal to the preset threshold (e.g., a preset threshold of 3). If not, it adds the default transform structure to the candidate transform structure list, performs validity checks and deduplication, and / or, if yes, it parses the list index encoding result to determine the candidate transform structure, and performs inverse quantization and inverse transform based on the candidate transform structure to obtain the residual block and prediction block, thereby reconstructing the current block output and / or reference buffer.

[0422] In this embodiment, multiple syntax elements, such as the traditional independent indicator sub-block transform state, inter-frame multiple transform selection index, and low-frequency inseparable transform index, are integrated into a unified transform structure identifier and list index encoding result for transmission. This simplifies the traditional process of dispersing and combining multiple syntax elements into determining the predefined complete transform structure pointed to by a single index and encoding result. Only the transform structure identifier and list index encoding result need to be transmitted to completely determine the transform mode of the current block, reducing the bit overhead of the bitstream used to describe complex transform modes and improving bitstream compression efficiency.

[0423] Fifth embodiment

[0424] This application also provides a processing device, please refer to... Figure 15 , Figure 15 This is a functional block diagram of the processing device of this application, which can be installed in or is the processing equipment. The processing device includes:

[0425] Processing module A10 is used to transform the current block according to at least one transformation structure.

[0426] Optionally, at least one transformation structure includes at least one transformation method.

[0427] Optionally, at least one transformation structure is determined or obtained based on at least one of the following:

[0428] The transformation structure of at least one sub-block of the current block;

[0429] List of candidate transformation structures;

[0430] The prediction pattern for the current block;

[0431] The current block size;

[0432] The size of at least one sub-block;

[0433] Bitstream.

[0434] Optionally, the processing apparatus further includes at least one of the following:

[0435] At least one of the following transformation methods is included: discrete cosine transform, transform skip, sub-block transform, low-frequency non-separable quadratic transform, non-separable master transform, and multiple transform selection;

[0436] The candidate transformation structure list includes: at least one candidate transformation structure and / or a list index of at least one candidate transformation structure.

[0437] Optionally, at least one candidate transformation structure is determined or obtained based on at least one of the following:

[0438] Default transformation structure;

[0439] The transformation structure of at least one of the following: the block above the current block, the block above the current block, the block to the left of the current block, the block to the left of the current block, the block to the upper left of the current block, and the block to the upper left of the current block;

[0440] Transformation structure of at least one of the following: the upper adjacent sub-block, the upper non-adjacent sub-block, the left adjacent sub-block, the left non-adjacent sub-block, the upper left adjacent sub-block, and the upper left non-adjacent sub-block;

[0441] Transformation structure of at least one of the following: cross-component block, same-component block, same-position block, and time-domain block of the current block;

[0442] Transformation structures of at least one sub-block across sub-blocks, same-component sub-blocks, co-position sub-blocks, and time-domain sub-blocks.

[0443] Optionally, the processing apparatus further includes at least one of the following:

[0444] At least one transformation structure is determined or obtained based on at least one candidate transformation structure in the candidate transformation structure list;

[0445] At least one transformation structure is determined or obtained based on the transformation signaling in the code stream;

[0446] Block dimensions include at least one of the following: width, height, aspect ratio, perimeter, and area;

[0447] The index of the list of at least one candidate transformation structure is determined or obtained based on the sorting method of the candidate transformation structure list and / or the sorting position of at least one candidate transformation structure.

[0448] Optionally, the sorting position of at least one candidate transformation structure is determined or obtained based on at least one of the following:

[0449] The current block size;

[0450] The size of at least one sub-block;

[0451] The block size of at least one of the following: the block above the current block, the block above the current block, the block to the left of the current block, the block to the left of the current block, the block to the left of the current block, the block to the left of the current block, the block across components, the block in the same component, the block in the same position, and the block in the time domain.

[0452] The block size of at least one of the following: the upper adjacent sub-block, the upper non-adjacent sub-block, the left adjacent sub-block, the left non-adjacent sub-block, the upper left adjacent sub-block, the upper left non-adjacent sub-block, the cross-component sub-block, the same component sub-block, the same position sub-block, and the time domain sub-block;

[0453] The reuse frequency of at least one candidate transformation structure.

[0454] Optionally, the transformation signaling is determined or obtained based on at least one of the following:

[0455] Transformation structure identifier;

[0456] List index encoding result;

[0457] The number of candidate transformation structures;

[0458] Syntax elements for transformation methods;

[0459] Transform the kernel encoding result.

[0460] Optionally, the processing apparatus further includes at least one of the following:

[0461] At least one transformation structure identifier is determined or obtained based on the number of candidate transformation structures;

[0462] At least one transformation structure identifier is determined or obtained based on the transformation cost;

[0463] At least one list index encoding result includes at least one of the following: a first unary code encoding result, a first truncated unary code encoding result, and a first exponential Columbus encoding result;

[0464] At least one transform kernel encoding result includes at least one of the following: second unary code encoding result, second truncated unary code encoding result, and second exponential Columbus encoding result.

[0465] The processing device provided in this application embodiment is similar in implementation principle and beneficial effect to the technical solution shown in the corresponding method embodiment above, and will not be described again here.

[0466] This application also provides a processing device, including a memory and a processor. The memory stores a processing program, and when the processing program is executed by the processor, it implements the steps of the processing method in any of the above embodiments.

[0467] This application also provides a storage medium storing a processing program, which, when executed by a processor, implements the steps of the processing method in any of the above embodiments.

[0468] In the embodiments of the processing device and storage medium provided in this application, all the technical features of any of the above-described processing method embodiments may be included. The extended and explanatory content of the specification is basically the same as that of the embodiments of the above methods, and will not be repeated here.

[0469] This application also provides a computer program product, which includes computer program code. When the computer program code is run on a computer, it causes the computer to perform the methods described in the various possible implementations above.

[0470] This application also provides a chip, including a memory and a processor. The memory is used to store a computer program, and the processor is used to call and run the computer program from the memory, so that a device with the chip installed performs the methods described in the various possible implementations above.

[0471] It is understood that the above scenarios are merely examples and do not constitute a limitation on the application scenarios of the technical solutions provided in the embodiments of this application. The technical solutions of this application can also be applied to other scenarios. For example, as those skilled in the art will know, with the evolution of system architecture and the emergence of new business scenarios, the technical solutions provided in the embodiments of this application are also applicable to similar technical problems.

[0472] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.

[0473] The steps in the method of this application embodiment can be adjusted, combined, or deleted according to actual needs.

[0474] The units in the device of this application embodiment can be merged, divided, and deleted according to actual needs.

[0475] In this application, the same or similar terms, concepts, technical solutions and / or application scenario descriptions are generally described in detail only when they appear for the first time. When they appear again, they are generally not repeated for the sake of brevity. When understanding the technical solutions and other contents of this application, the same or similar terms, concepts, technical solutions and / or application scenario descriptions that are not described in detail later can be referred to their previous relevant detailed descriptions.

[0476] In this application, the descriptions of the various embodiments have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0477] The technical features of the present application can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of the present application.

[0478] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) as described above, and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, controlled terminal, or network device, etc.) to execute the methods of each embodiment of this application.

[0479] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product. A computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the flow or function according to the embodiments of this application is generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a storage medium or transmitted from one storage medium to another. For example, computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) means. The storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, storage disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid state disk (SSD)).

[0480] The above are merely preferred embodiments of this application and do not limit the patent scope of this application. Any equivalent structural or procedural transformations made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of this application.

Claims

1. A processing method, characterized in that, Including the following steps: S10, Transform the current block according to at least one transformation structure; The transformation structure includes a transformation method for performing residual transformation on the current block, and the enabling status and segmentation method information of sub-block transformation; At least one transform structure is determined or obtained based on at least one candidate transform structure in the candidate transform structure list and the transform signaling in the code stream; The candidate transformation structure list includes: at least one candidate transformation structure and a list index of at least one candidate transformation structure; At least one candidate transformation structure is determined or obtained based on at least one of the transformation structures of the current block's upper adjacent block, upper non-adjacent block, left adjacent block, left non-adjacent block, upper left adjacent block, and upper left non-adjacent block, and / or at least one of the transformation structures of the current block's upper adjacent sub-block, upper non-adjacent sub-block, left adjacent sub-block, left non-adjacent sub-block, upper left adjacent sub-block, and upper left non-adjacent sub-block: The index of the list of at least one candidate transformation structure is determined or obtained based on the sorting method of the candidate transformation structure list and / or the sorting position of at least one candidate transformation structure; The transformation signaling is determined or obtained based on the list index encoding result.

2. The processing method as described in claim 1, characterized in that, At least one transformation structure can also be determined or obtained based on at least one of the following: The transformation structure of at least one sub-block of the current block; The prediction pattern for the current block; The current block size; The block size of at least one sub-block.

3. The processing method as described in claim 2, characterized in that, Block dimensions include at least one of the following: width, height, aspect ratio, perimeter, and area.

4. The processing method as described in claim 1, characterized in that, At least one of the following transformation methods is included: discrete cosine transform, transform skip, low-frequency non-separable quadratic transform, non-separable master transform, and multiple transform selection.

5. The processing method as described in claim 1, characterized in that, At least one candidate transformation structure is determined or obtained based on at least one of the following: Default transformation structure; Transformation structure of at least one of the following: cross-component block, same-component block, same-position block, and time-domain block of the current block; Transformation structures of at least one sub-block across sub-blocks, same-component sub-blocks, co-position sub-blocks, and time-domain sub-blocks.

6. The processing method as described in claim 1, characterized in that, The sorting position of at least one candidate transformation structure is determined or obtained based on at least one of the following: The current block size; The size of at least one sub-block; The block size of at least one of the following: the block above the current block, the block above the current block, the block to the left of the current block, the block to the left of the current block, the block to the left of the current block, the block to the left of the current block, the block across components, the block in the same component, the block in the same position, and the block in the time domain. The block size of at least one of the following: the upper adjacent sub-block, the upper non-adjacent sub-block, the left adjacent sub-block, the left non-adjacent sub-block, the upper left adjacent sub-block, the upper left non-adjacent sub-block, the cross-component sub-block, the same component sub-block, the same position sub-block, and the time domain sub-block; The reuse frequency of at least one candidate transformation structure.

7. The processing method as described in claim 1, characterized in that, The transformation signaling can also be determined or obtained based on at least one of the following: Transformation structure identifier; The number of candidate transformation structures; Syntax elements for transformation methods; Transform the kernel index encoding result.

8. The processing method as described in claim 7, characterized in that, The method further includes at least one of the following: At least one transformation structure identifier is determined or obtained based on the number of candidate transformation structures; At least one list index encoding result includes at least one of the following: a first unary code encoding result, a first truncated unary code encoding result, and a first exponential Columbus encoding result; At least one transform kernel index encoding result includes at least one of the following: second unary code encoding result, second truncated unary code encoding result, and second exponential Columbus encoding result.

9. A processing device, characterized in that, include: A memory and a processor, wherein the memory stores a processing program, and the processing program, when executed by the processor, implements the steps of the processing method as described in any one of claims 1 to 8.

10. A storage medium, characterized in that, The storage medium stores a computer program, which, when executed by a processor, implements the steps of the processing method as described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Encoding and decoding method, encoder, decoder and storage medium

    CN119487848A