Storing and signaling post-processing order in a header
By introducing a signaling mechanism and storing the order of post-processing filters in the video coding system, especially the order of luminance and chrominance post-filters, the problem of poor filter order storage and application in the prior art is solved, thereby improving coding efficiency and image quality.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-10
- Publication Date
- 2026-03-27
AI Technical Summary
Existing video coding systems lack effective signaling mechanisms for storing and applying post-processing filter order, making it difficult to optimize coding efficiency and quality.
Introducing signaling mechanisms into video encoding and decoding devices involves storing the post-processing filter order, particularly the luma and chroma post-filter order, in the header of a region or image. The filter order is determined based on encoding cost and distortion optimization, and filters such as Cross Component Sample Adaptive Offset (CCSAO) are applied before SAO.
It improves the coding efficiency and image quality of video encoding, optimizes coding cost and distortion by dynamically adjusting the filter order, and enhances the processing efficiency and image reconstruction effect of encoder and decoder.
Smart Images

Figure CN121753327A_ABST
Abstract
Description
[0001] Cross Reference to Related Applications This application claims the benefit of European Provisional Patent Application No. 23306072.2, filed June 29, 2023, the contents of which are incorporated herein by reference. BACKGROUND
[0002] Video coding systems can be used to compress digital video signals, for example, to reduce the storage and / or transmission bandwidth required for such signals. Video coding systems can include, for example, block-based, wavelet-based, and / or object-based systems. SUMMARY
[0003] Systems, methods, and instrumentalities are disclosed for storing post-processing order in a region or picture header with a dedicated code. Signaling can be used to define a post filter order. Signaling can specify an order for post-processing at an encoder and a decoder. The post-processing order can be stored in a header. In an example, the post-processing order can be stored in a header associated with a region or slice, or stored in a picture header.
[0004] In an example, a video coding device can determine a post filter order associated with a region (or slice) or a sub-region or coding tree unit (CTU) of a picture. The post filter order can indicate a post-processing order. In one example, cross component sample adaptive offset (CCSAO) can be applied before sample adaptive offset (SAO). The post filter order can be a luma post filter order or a chroma post filter order.
[0005] The video coding device can send the post filter order in a region (or slice) header or a picture header. The video coding device can determine a luma post filter order and apply the luma post filter order to a chroma post filter order. The luma post filter order can be different from the chroma post filter order.
[0006] The video coding device can determine the post filter order based on coding cost and distortion (e.g., optimization of coding cost and distortion). When sending the post filter order in a region header, the video coding device can determine a rate-distortion cost for each sub-region of a region. When sending the post filter order in a picture header, the video coding device can determine a rate-distortion cost for each sub-region of a picture. The video coding device can determine a rate-distortion cost based on a distortion of a current block with its coding parameters, an associated rate or cost, and / or a parameter derived from a quantization parameter.
[0007] Systems, methods, and instrumentalities are disclosed for video encoding. A video encoding device can include a processor configured to determine a post filter order associated with a region or a sub-region of a picture. The post filter order can be associated with at least two filters. The post filter order can indicate a post filter processing order. The video encoding device can signal the post filter order in a region header or a picture header.
[0008] A video encoding device as described herein can include one or more features. In an example, the video encoding device can be configured to generate an encoding cost and a distortion. The encoding cost and the distortion can be associated with the post filter processing order. The video encoding device can be configured to determine a luma post filter order or a chroma post filter order. The video encoding device can be configured to determine a luma post filter order. The video encoding device can be configured to apply a luma post filter order to a chroma post filter order. In an example, the luma post filter order can be different than the chroma post filter order. The video encoding device can be configured to determine a rate-distortion cost of a sub-region of the region when signaling the post filter order to the region header.
[0009] The video encoding device can be configured to determine a rate-distortion cost of a sub-region of the picture when signaling the post filter order to the picture header. The video encoding device can be configured to determine the rate-distortion cost based on a distortion of a current block having an encoding parameter of the current block, an associated rate or cost, and / or a parameter derived from a quantization parameter. In an example, the sub-region can be a coding tree unit (CTU), the region can be a slice, and the region header can be a slice header. The at least two filters can include a sample adaptive offset (SAO), a cross component SAO (CCSAO), an adaptive loop filter, or a deblocking filter. The post filter order can indicate that CCSAO is to be applied before SAO.
[0010] Systems, methods, and instrumentalities are disclosed for video decoding. A video decoding device can include a processor configured to determine a post filter order associated with a region or a sub-region of a picture. The post filter order can be associated with at least two filters. The post filter order can indicate a post filter processing order. The video decoding device can signal the post filter order in a region header or a picture header.
[0011] The video decoding apparatus described herein may include one or more features. In an example, the video decoding apparatus may be configured to generate encoding costs and distortion. The encoding costs and distortion may be associated with the post-filter processing order. The video decoding apparatus may be configured to determine a luma post-filter order or a chroma post-filter order. The video decoding apparatus may be configured to determine a luma post-filter order. The video decoding apparatus may be configured to apply the luma post-filter order to the chroma post-filter order. In an example, the luma post-filter order may be different from the chroma post-filter order. The video decoding apparatus may be configured to determine the rate distortion cost of a sub-region of the region when the post-filter order is sent to the region header.
[0012] The video decoding device can be configured to determine the rate distortion cost of a sub-region of the image when the post-filter order is sent to the image header. The video decoding device can be configured to determine the rate distortion cost based on the distortion of the current block with the coding parameters of the current block, the associated rate or cost, and / or parameters derived from the quantization parameters. In an example, the sub-region may be a coding tree unit (CTU), the region may be a slice, and the region header may be a slice header. The at least two filters may include a Sample Adaptive Offset (SAO), a Cross Component SAO (CCSAO), an Adaptive Loop Filter, or a Deblocking Filter. The post-filter order may indicate that CCSAO should be applied before the SAO.
[0013] The systems, methods, and means described herein may relate to a decoder. In some examples, the systems, methods, and means described herein may relate to an encoder. In some examples, the systems, methods, and means described herein may relate to signals (e.g., from an encoder and / or received by a decoder). A computer-readable medium may include instructions for causing one or more processors to perform the methods described herein. A computer program product may include instructions that, when executed by one or more processors, cause the one or more processors to perform the methods described herein. Attached Figure Description
[0014] FIG. 1A This is a system diagram illustrating an example communication system in which one or more of the disclosed embodiments may be implemented.
[0015] FIG. 1B The illustration shows a method according to one embodiment. FIG. 1A The diagram shows a system diagram of an example wireless transmit / receive unit (WTRU) used in a communication system.
[0016] FIG. 1CThe illustration shows a method according to one embodiment. FIG. 1A The diagram illustrates a system diagram of an example radio access network (RAN) and an example core network (CN) used in the communication system.
[0017] FIG. 1D The illustration shows a method according to one embodiment. FIG. 1A The illustrated system diagram shows yet another example RAN and yet another example CN used in the communication system.
[0018] FIG. 2 The illustration shows an example video encoder.
[0019] FIG. 3 The illustration shows an example video decoder.
[0020] FIG. 4 The illustration shows an example of a system that can implement various aspects and examples.
[0021] FIG. 5A-5C The illustration shows an example of reconstructing the sample category in the case of edge offset (EO) mode.
[0022] FIG. 6 The illustration shows an example of a pixel range uniformly divided into multiple frequency bands, for example, in band offset (BO) mode.
[0023] FIG. 7 The illustration shows an example of the modified Sample Adaptive Offset (SAO) process when Cross Component Sample Adaptive Offset (CCSAO) is applied.
[0024] FIG. 8 An example of candidate locations for the CCSAO classifier is illustrated.
[0025] FIG. 9 The illustration shows an example of joint clipping after adding SAO / Blank Filter (BIF) / CCSAO offset to the input sample.
[0026] FIG. 10 Examples of four one-dimensional (1-D) orientation patterns for CCSAO EO sample classification are illustrated: horizontal, vertical, 135° diagonal, and 45° diagonal.
[0027] FIG. 11 An example of the filter shape of an adaptive loop filter (ALF) is illustrated.
[0028] FIG. 12A-12D The illustration shows an example of a subsampling location used for subsampling Laplace calculation.
[0029] FIG. 13AThe illustration shows an example of the placement of the CC-ALF relative to other loop filters.
[0030] FIG. 13B An example of a diamond-shaped filter is illustrated.
[0031] FIG. 14 An example of the post-processing filter order is illustrated.
[0032] FIG. 15 An example of the post-processing filter order is illustrated. Detailed Implementation
[0033] A more detailed understanding can be obtained from the description given below with reference to the accompanying drawings and examples.
[0034] FIG. 1A This is a schematic diagram illustrating an example communication system 100 in which one or more of the disclosed embodiments may be implemented. The communication system 100 may be a multi-access system that provides content such as voice, data, video, messages, and broadcasts to multiple wireless users. The communication system 100 enables multiple wireless users to access such content by sharing system resources, including wireless bandwidth. For example, the communication system 100 may employ one or more channel access methods, such as Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Frequency Division Multiple Access (FDMA), Orthogonal FDMA (OFDMA), Single Carrier FDMA (SC-FDMA), Zero-Tail Unique Word DFT Spread Spectrum OFDM (ZT UWDTS-s OFDM), Unique Word OFDM (UW-OFDM), Resource Block Filtered OFDM, Filter Bank Multicarrier (FBMC), etc.
[0035] like FIG. 1AAs shown, the communication system 100 may include wireless transmit / receive units (WTRUs) 102a, 102b, 102c, 102d, RAN 104 / 113, CN 106 / 115, Public Switched Telephone Network (PSTN) 108, Internet 110, and other networks 112. However, it should be understood that the disclosed embodiments contemplate any number of WTRUs, base stations, networks, and / or network elements. Each of the WTRUs 102a, 102b, 102c, and 102d may be any type of device configured to operate and / or communicate in a wireless environment. For example, WTRUs 102a, 102b, 102c, and 102d (any of which may be referred to as a “station” and / or “STA”) may be configured to transmit and / or receive wireless signals and may include user equipment (UE), mobile stations, fixed or mobile subscriber units, subscription-based units, pagers, cellular phones, personal digital assistants (PDAs), smartphones, laptops, netbooks, personal computers, wireless sensors, hotspots or Mi-Fi devices, Internet of Things (IoT) devices, watches or other wearable devices, head-mounted displays (HMDs), vehicles, drones, medical devices and applications (e.g., remote surgery), industrial devices and applications (e.g., robots and / or other wireless devices operating in industrial and / or automated processing chain environments), consumer electronics devices, devices operating on commercial and / or industrial wireless networks, etc. Any of WTRUs 102a, 102b, 102c, and 102d may be interchangeably referred to as a UE.
[0036] The communication system 100 may also include base station 114a and / or base station 114b. Each of base stations 114a and 114b may be any type of device configured to wirelessly interface with at least one of WTRUs 102a, 102b, 102c, and 102d to facilitate access to one or more communication networks, such as CN 106 / 115, the Internet 110, and / or other networks 112. For example, base stations 114a and 114b may be base transceiver stations (BTS), node B, eNode B, home node B, home eNode B, gNB, NR node B, site controller, access point (AP), wireless router, etc. Although base stations 114a and 114b are each depicted as a single element, it should be understood that base stations 114a and 114b may include any number of interconnected base stations and / or network elements.
[0037] Base station 114a may be part of RAN 104 / 113, which may also include other base stations and / or network elements (not shown), such as base station controllers (BSCs), radio network controllers (RNCs), relay nodes, etc. Base station 114a and / or base station 114b may be configured to transmit and / or receive radio signals on one or more carrier frequencies, which may be referred to as cells (not shown). These frequencies may be in licensed spectrum, unlicensed spectrum, or a combination of licensed and unlicensed spectrum. A cell may provide coverage of a specific geographic area, which may be relatively fixed or may change over time. The cell may be further divided into cell sectors. For example, the cell associated with base station 114a may be divided into three sectors. Therefore, in one embodiment, base station 114a may include three transceivers, i.e., one transceiver for each sector of the cell. In one embodiment, base station 114a may employ multiple-input multiple-output (MIMO) technology and may use multiple transceivers for each sector of the cell. For example, beamforming can be used to transmit and / or receive signals in a desired spatial direction.
[0038] Base stations 114a and 114b can communicate with one or more of WTRUs 102a, 102b, 102c, and 102d via air interface 116. Air interface 116 can be any suitable wireless communication link (e.g., radio frequency (RF), microwave, centimeter wave, micrometer wave, infrared (IR), ultraviolet (UV), visible light, etc.). Any suitable radio access technology (RAT) can be used to establish air interface 116.
[0039] More specifically, as described above, the communication system 100 can be a multi-access system and can employ one or more channel access schemes, such as CDMA, TDMA, FDMA, OFDMA, SC-FDMA, etc. For example, base stations 114a and WTRUs 102a, 102b, and 102c in RAN 104 / 113 can implement radio technologies such as Universal Mobile Telecommunications System (UMTS) Terrestrial Radio Access (UTRA), which can establish air interfaces 115 / 116 / 117 using Wideband CDMA (WCDMA). WCDMA may include communication protocols such as High-Speed Packet Access (HSPA) and / or evolved HSPA (HSPA+). HSPA may include High-Speed Downlink (DL) Packet Access (HSDPA) and / or High-Speed UL Packet Access (HSUPA).
[0040] In one embodiment, base station 114a and WTRUs 102a, 102b, 102c may implement radio technologies such as evolved UMTS terrestrial radio access (E-UTRA), which may use Long Term Evolution (LTE) and / or Advanced LTE (LTE-A) and / or Advanced LTE Pro (LTE-A Pro) to establish air interface 116.
[0041] In one embodiment, base station 114a and WTRUs 102a, 102b, 102c can implement radio technologies such as NR radio access, which can establish an air interface 116 using a new radio (NR).
[0042] In one embodiment, base station 114a and WTRUs 102a, 102b, and 102c can implement multiple radio access technologies. For example, base station 114a and WTRUs 102a, 102b, and 102c can jointly implement LTE radio access and NR radio access, for example, using the dual connectivity (DC) principle. Therefore, the air interface used by WTRUs 102a, 102b, and 102c can be characterized by multiple types of radio access technologies and / or transmissions sent to / from multiple types of base stations (e.g., eNBs and gNBs).
[0043] In other embodiments, base station 114a and WTRUs 102a, 102b, 102c can implement radio technologies such as IEEE 802.11 (i.e., Wi-Fi), IEEE 802.16 (i.e., WiMAX), CDMA2000, CDMA2000 1X, CDMA2000 EV-DO, Provisional Standard 2000 (IS-2000), Provisional Standard 95 (IS-95), Provisional Standard 856 (IS-856), Global System for Mobile Communications (GSM), Enhanced Data Rate GSM Evolution (EDGE), GSM EDGE (GERAN), etc.
[0044] For example, FIG. 1ABase station 114b can be a wireless router, home node B, home eNodeB, or access point, and can utilize any suitable RAT to facilitate wireless connectivity in a local area, such as commercial locations, homes, vehicles, campuses, industrial facilities, air corridors (e.g., for drone use), roads, etc. In one embodiment, base station 114b and WTRUs 102c, 102d can implement radio technologies such as IEEE 802.11 to establish a wireless local area network (WLAN). In one embodiment, base station 114b and WTRUs 102c, 102d can implement radio technologies such as IEEE 802.15 to establish a wireless personal area network (WPAN). In yet another embodiment, base station 114b and WTRUs 102c, 102d can utilize cellular-based RATs (e.g., WCDMA, CDMA2000, GSM, LTE, LTE-A, LTE-A Pro, NR, etc.) to establish picocells or femtocells. FIG. 1A As shown, base station 114b can be directly connected to the Internet 110. Therefore, base station 114b may not need to access the Internet 110 via CN 106 / 115.
[0045] RAN 104 / 113 can communicate with CN 106 / 115, which can be any type of network configured to provide voice, data, application, and / or Voice over Internet Protocol (VoIP) services to one or more of WTRUs 102a, 102b, 102c, and 102d. Data can have different Quality of Service (QoS) requirements, such as different throughput requirements, latency requirements, fault tolerance requirements, reliability requirements, data throughput requirements, mobility requirements, etc. CN 106 / 115 can provide call control, billing services, location-based services, prepaid calling, internet connectivity, video distribution, and / or perform advanced security functions such as user authentication. Although in FIG. 1A Although not shown, it should be understood that RAN104 / 113 and / or CN 106 / 115 can communicate directly or indirectly with other RANs that use the same RAT as or a different RAT than RAN 104 / 113. For example, in addition to being connected to RAN 104 / 113, which may utilize NR radio technology, CN 106 / 115 can also communicate with another RAN (not shown) that uses GSM, UMTS, CDMA 2000, WiMAX, E-UTRA, or WiFi radio technology.
[0046] CN 106 / 115 can also serve as a gateway for WTRU 102a, 102b, 102c, 102d to access PSTN 108, the Internet 110, and / or other networks 112. PSTN 108 may include a circuit-switched telephone network providing Common Old-Style Telephone Service (POTS). The Internet 110 may include a global system of interconnected computer networks and devices using common communication protocols such as Transmission Control Protocol (TCP), User Datagram Protocol (UDP), and / or Internet Protocol (IP) from the TCP / IP Internet Protocol suite. Network 112 may include wired and / or wireless communication networks owned and / or operated by other service providers. For example, network 112 may include another CN connected to one or more RANs, which may use the same RAT as RAN 104 / 113 or a different RAT.
[0047] Some or all of the WTRUs 102a, 102b, 102c, and 102d in the communication system 100 may include multi-mode capabilities (e.g., WTRUs 102a, 102b, 102c, and 102d may include multiple transceivers for communicating with different wireless networks via different wireless links). For example... FIG. 1A The WTRU 102c shown can be configured to communicate with base station 114a, which may employ cellular-based radio technology, and to communicate with base station 114b, which may employ IEEE 802 radio technology.
[0048] FIG. 1B This is a system diagram illustrating example WTRU 102. (Example:) FIG. 1B As shown, among other things, WTRU 102 may include, in particular, a processor 118, a transceiver 120, a transmit / receive element 122, a speaker / microphone 124, a keypad 126, a display / touchpad 128, non-removable memory 130, removable memory 132, a power supply 134, a global positioning system (GPS) chipset 136, and / or other peripheral devices 138, etc. It should be understood that WTRU 102 may include any sub-combination of the foregoing elements while remaining consistent with the embodiments.
[0049] Processor 118 may be a general-purpose processor, a special-purpose processor, a conventional processor, a digital signal processor (DSP), multiple microprocessors, one or more microprocessors associated with a DSP core, a controller, a microcontroller, an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) circuit, any other type of integrated circuit (IC), a state machine, etc. Processor 118 may perform signal encoding, data processing, power control, input / output processing, and / or any other functions that enable WTRU 102 to operate in a wireless environment. Processor 118 may be coupled to transceiver 120, which may be coupled to transmitting / receiving element 122. Although FIG. 1B The processor 118 and transceiver 120 are depicted as separate components, but it should be understood that the processor 118 and transceiver 120 may be integrated together in an electronic package or chip.
[0050] Transmitting / receiving element 122 can be configured to transmit signals to or receive signals from a base station (e.g., base station 114a) over air interface 116. For example, in one embodiment, transmitting / receiving element 122 can be an antenna configured to transmit and / or receive RF signals. In one embodiment, transmitting / receiving element 122 can be, for example, a transmitter / detector configured to transmit and / or receive IR, UV, or visible light signals. In yet another embodiment, transmitting / receiving element 122 can be configured to transmit and / or receive both RF and optical signals. It should be understood that transmitting / receiving element 122 can be configured to transmit and / or receive any combination of wireless signals.
[0051] Although the transmitting / receiving element 122 is in FIG. 1B While depicted as a single element, WTRU 102 may include any number of transmit / receive elements 122. More specifically, WTRU 102 may employ MIMO technology. Thus, in one embodiment, WTRU 102 may include two or more transmit / receive elements 122 (e.g., multiple antennas) for transmitting and receiving wireless signals on air interface 116.
[0052] Transceiver 120 can be configured to modulate signals transmitted by transmitting / receiving element 122 and demodulate signals received by transmitting / receiving element 122. As described above, WTRU 102 can have multi-mode capability. Therefore, for example, transceiver 120 may include multiple transceivers to enable WTRU 102 to communicate via multiple RATs, such as NR and IEEE 802.11.
[0053] The processor 118 of WTRU 102 can be coupled to a speaker / microphone 124, a keypad 126, and / or a display / touchpad 128 (e.g., a liquid crystal display (LCD) unit or an organic light-emitting diode (OLED) display unit) and can receive user input data therefrom. The processor 118 can also output user data to the speaker / microphone 124, keypad 126, and / or display / touchpad 128. Furthermore, the processor 118 can access and store information from any type of suitable memory, such as non-removable memory 130 and / or removable memory 132. Non-removable memory 130 may include random access memory (RAM), read-only memory (ROM), a hard disk, or any other type of memory storage device. Removable memory 132 may include a user identification module (SIM) card, memory stick, secure digital storage (SD) card, etc. In other embodiments, the processor 118 can access and store information from memory that is not physically located on WTRU 102 (e.g., a server or home computer (not shown)).
[0054] The processor 118 can receive power from the power supply 134 and can be configured to distribute and / or control power to other components in the WTRU 102. The power supply 134 can be any suitable device that powers the WTRU 102. For example, the power supply 134 may include one or more dry cell batteries (e.g., nickel-cadmium (NiCd), nickel-zinc (NiZn), nickel metal hydride (NiMH), lithium-ion (Li-ion), etc.), solar cells, fuel cells, etc.
[0055] The processor 118 may also be coupled to a GPS chipset 136, which may be configured to provide location information (e.g., longitude and latitude) about the current location of the WTRU 102. In addition to, or instead of, information from the GPS chipset 136, the WTRU 102 may receive location information on the air interface 116 from base stations (e.g., base stations 114a, 114b) and / or determine its location based on the timing of signals received from two or more nearby base stations. It should be understood that the WTRU 102 may acquire location information using any suitable location determination method while remaining consistent with the embodiments.
[0056] The processor 118 may be further coupled to other peripheral devices 138, which may include one or more software and / or hardware modules providing additional features, functions, and / or wired or wireless connectivity. For example, peripheral devices 138 may include accelerometers, electronic compasses, satellite transceivers, digital cameras (for photos and / or videos), Universal Serial Bus (USB) ports, vibration devices, television transceivers, hands-free headsets, Bluetooth® modules, FM radio units, digital music players, media players, video game player modules, internet browsers, virtual reality and / or augmented reality (VR / AR) devices, activity trackers, etc. Peripheral devices 138 may include one or more sensors, which may be one or more of the following: gyroscopes, accelerometers, Hall effect sensors, magnetometers, orientation sensors, proximity sensors, temperature sensors, time sensors; geolocation sensors, altimeters, light sensors, touch sensors, magnetometers, barometers, attitude sensors, biosensors, and / or humidity sensors.
[0057] WTRU 102 may include a full-duplex radio for which the transmission and reception of some or all signals (e.g., signals associated with specific subframes for UL (e.g., for transmission) and downlink (e.g., for reception)) may be concurrent and / or simultaneous. The full-duplex radio may include an interference management unit to reduce and / or substantially eliminate self-interference via hardware (e.g., chokes) or via signal processing by a processor (e.g., a separate processor (not shown) or via processor 118). In one embodiment, WTRU 102 may include a half-duplex radio for which the transmission and reception of some or all signals (e.g., signals associated with specific subframes for UL (e.g., for transmission) or downlink (e.g., for reception)) may be concurrent and / or simultaneous.
[0058] FIG. 1C This diagram illustrates a system diagram of RAN 104 and CN 106 according to an embodiment. As described above, RAN 104 can communicate with WTRUs 102a, 102b, and 102c via air interface 116 using E-UTRA radio technology. RAN 104 can also communicate with CN 106.
[0059] RAN 104 may include eNode-Bs 160a, 160b, and 160c; however, it should be understood that RAN 104 may include any number of eNode-Bs while remaining consistent with the embodiments. eNode-Bs 160a, 160b, and 160c may each include one or more transceivers for communicating with WTRUs 102a, 102b, and 102c on air interface 116. In one embodiment, eNode-Bs 160a, 160b, and 160c may implement MIMO technology. Therefore, for example, eNode-B 160a may use multiple antennas to transmit and / or receive radio signals from WTRU 102a.
[0060] Each of the eNode-B 160a, 160b, and 160c can be associated with a specific cell (not shown) and can be configured to handle radio resource management decisions, handover decisions, user scheduling in UL and / or DL, etc. FIG. 1C As shown, eNode-B 160a, 160b, and 160c can communicate with each other on the X2 interface.
[0061] FIG. 1C The CN 106 shown may include a Mobility Management Entity (MME) 162, a Serving Gateway (SGW) 164, and a Packet Data Network (PDN) Gateway (or PGW) 166. While each of the foregoing elements is described as part of CN 106, it should be understood that any of these elements may be owned and / or operated by an entity other than a CN operator.
[0062] The MME 162 can connect to each of the eNode-Bs 162a, 162b, and 162c in RAN 104 via the S1 interface and can act as a control node. For example, the MME 162 can be responsible for authenticating users of WTRUs 102a, 102b, and 102c, bearer activation / deactivation, selecting a specific serving gateway during the initial attachment of WTRUs 102a, 102b, and 102c, etc. The MME 162 can provide control plane functions for handover between RAN 104 and other RANs (not shown) employing other radio technologies such as GSM and / or WCDMA.
[0063] The SGW 164 can connect to each of the eNode Bs 160a, 160b, and 160c in RAN 104 via the S1 interface. The SGW 164 can typically route and forward user data packets to / from WTRUs 102a, 102b, and 102c. The SGW 164 can perform other functions, such as anchoring the user plane during inter-eNode B handover, triggering paging when DL data is available for WTRUs 102a, 102b, and 102c, and managing and storing the context of WTRUs 102a, 102b, and 102c.
[0064] SGW 164 can connect to PGW 166, which can provide WTRU 102a, 102b, 102c with access to packet-switched networks such as Internet 110, so as to facilitate communication between WTRU 102a, 102b, 102c and IP-enabled devices.
[0065] CN 106 can facilitate communication with other networks. For example, CN 106 can provide WTRU 102a, 102b, and 102c with access to a circuit-switched network such as PSTN 108, facilitating communication between WTRU 102a, 102b, and 102c and traditional landline communication equipment. For example, CN 106 may include, or be able to communicate with, an IP gateway (e.g., an IP Multimedia Subsystem (IMS) server) that acts as an interface between CN 106 and PSTN 108. Furthermore, CN 106 can provide WTRU 102a, 102b, and 102c with access to other networks 112, which may include other wired and / or wireless networks owned and / or operated by other service providers.
[0066] Despite WTRU in FIG. 1A-1D While described as a wireless terminal, it is conceivable that, in some representative embodiments, such a terminal may use (e.g., temporarily or permanently) a wired communication interface with a communication network.
[0067] In a representative embodiment, another network 112 may be a WLAN.
[0068] A WLAN in Infrastructure Basic Services Set (BSS) mode can have an access point (AP) for the BSS and one or more stations (STAs) associated with the AP. The AP can access or interface with a distributed system (DS) or another type of wired / wireless network that transmits traffic to and / or out of the BSS. Traffic originating outside the BSS destined for a STA can reach and be delivered to the STA via the AP. Traffic originating from a STA destined for an external BSS can be sent to the AP for delivery to the appropriate destination. For example, traffic between STAs within the BSS can be transmitted via the AP, where the source STA can send traffic to the AP, and the AP can deliver traffic to the destination STA. Traffic between STAs within the BSS can be considered and / or referred to as peer-to-peer traffic. Peer-to-peer traffic can be transmitted between source and destination STAs (e.g., directly between them) using Direct Link Establishment (DLS). In some representative embodiments, the DLS can use 802.11e DLS or 802.11z Tunneled DLS (TDLS). A WLAN using the Standalone BSS (IBSS) mode may not have an access point (AP), and STAs within the IBSS or using the IBSS (e.g., all STAs) can communicate directly with each other. The IBSS communication mode is sometimes referred to here as an "ad-hoc" communication mode.
[0069] When using 802.11ac infrastructure operating mode or a similar operating mode, the AP can transmit beacons on a fixed channel, such as the primary channel. The primary channel can be of a fixed width (e.g., a wide bandwidth of 20 MHz) or dynamically set via signaling. The primary channel can be the operating channel of the BSS and can be used by the STA to establish a connection with the AP. In some representative embodiments, such as in an 802.11 system, Carrier Sense Multiple Access (CSMA / CA) with collision avoidance can be implemented. For CSMA / CA, each STA, including the AP, can sense the primary channel. If a particular STA senses / detects and / or determines that the primary channel is busy, that particular STA can back off. A single STA (e.g., only one station) can transmit at any given time within a given BSS.
[0070] High-throughput (HT) STAs can communicate using a 40 MHz wide channel, for example, by combining a primary 20 MHz channel with adjacent or non-adjacent 20 MHz channels.
[0071] Very High Throughput (VHT) STAs can support channels with widths of 20 MHz, 40 MHz, 80 MHz, and / or 160 MHz. 40 MHz and / or 80 MHz channels can be formed by combining consecutive 20 MHz channels. A 160 MHz channel can be formed by combining eight consecutive 20 MHz channels, or by combining two non-consecutive 80 MHz channels, which can be referred to as an 80+80 configuration. For the 80+80 configuration, after channel coding, the data passes through a segment resolver, which splits the data into two streams. Each stream can be processed separately using Inverse Fast Fourier Transform (IFFT) and time-domain processing. These streams can be mapped onto two 80 MHz channels, and the data can be transmitted by the transmitting STA. At the receiver of the receiving STA, the operation of the 80+80 configuration can be reversed, and the combined data can be sent to the Media Access Control (MAC).
[0072] 802.11af and 802.11ah support operating modes below 1 GHz. The channel operating bandwidth and carrier in 802.11af and 802.11ah are reduced compared to those used in 802.11n and 802.11ac. 802.11af supports 5 MHz, 10 MHz, and 20 MHz bandwidths in the TV whitespace (TVWS) spectrum, while 802.11ah supports 1 MHz, 2 MHz, 4 MHz, 8 MHz, and 16 MHz bandwidths using non-TVWS spectrum. According to a representative embodiment, 802.11ah can support metering-type control / machine-type communications, such as MTC devices in macro coverage areas. MTC devices may have certain capabilities, such as limited capabilities, including support for (e.g., only) certain and / or limited bandwidths. MTC devices may include batteries with a battery life exceeding a threshold (e.g., to maintain a very long battery life).
[0073] WLAN systems that can support multiple channels and channel bandwidths, such as 802.11n, 802.11ac, 802.11af, and 802.11ah, include channels that can be designated as the primary channel. The bandwidth of the primary channel can be equal to the maximum common operating bandwidth supported by all STAs in the BSS. The bandwidth of the primary channel can be set and / or limited by the STA among all STAs operating in the BSS that supports the minimum bandwidth operating mode. In the example of 802.11ah, for STAs that support (e.g., only support) the 1 MHz mode (e.g., MTC type devices), the primary channel can be 1 MHz wide, even if the AP and other STAs in the BSS support 2 MHz, 4 MHz, 8 MHz, 16 MHz, and / or other channel bandwidth operating modes. Carrier Sense and / or Network Allocation Vector (NAV) settings may depend on the status of the primary channel. If the primary channel is busy, for example, because an STA (which only supports the 1 MHz operating mode) is transmitting to the AP, the entire available band can be considered busy, even if most of the available band remains idle and can be available.
[0074] In the United States, the available frequency band for 802.11ah is from 902 MHz to 928 MHz. In South Korea, the available frequency band is from 917.5 MHz to 923.5 MHz. In Japan, the available frequency band is from 916.5 MHz to 927.5 MHz. The total available bandwidth for 802.11ah is 6 MHz to 26 MHz, depending on the country code.
[0075] FIG. 1D This diagram illustrates a system diagram of RAN 113 and CN 115 according to one embodiment. As described above, RAN 113 can communicate with WTRUs 102a, 102b, and 102c via air interface 116 using NR radio technology. RAN 113 can also communicate with CN 115.
[0076] RAN 113 may include gNBs 180a, 180b, and 180c; however, it should be understood that RAN 113 may include any number of gNBs while remaining consistent with the embodiments. gNBs 180a, 180b, and 180c may each include one or more transceivers for communicating with WTRUs 102a, 102b, and 102c on air interface 116. In one embodiment, gNBs 180a, 180b, and 180c may implement MIMO technology. For example, gNBs 180a and 180b may utilize beamforming to transmit signals to and / or receive signals from gNBs 180a, 180b, and 180c. Therefore, for example, gNB 180a may use multiple antennas to transmit radio signals to and / or receive radio signals from WTRU 102a. In one embodiment, gNBs 180a, 180b, and 180c can implement carrier aggregation technology. For example, gNB 180a can transmit multiple component carriers (not shown) to WTRU 102a. A subset of these component carriers can be on unlicensed spectrum, while the remaining component carriers can be on licensed spectrum. In one embodiment, gNBs 180a, 180b, and 180c can implement Coordinated Multipoint (CoMP) technology. For example, WTRU 102a can receive coordinated transmissions from gNBs 180a and 180b (and / or gNB 180c).
[0077] WTRUs 102a, 102b, and 102c can communicate with gNBs 180a, 180b, and 180c using transmissions associated with scalable digitization. For example, the OFDM symbol spacing and / or OFDM subcarrier spacing can differ for different transmissions, different cells, and / or different portions of the radio transmission spectrum. WTRUs 102a, 102b, and 102c can communicate with gNBs 180a, 180b, and 180c using subframes or transmission time intervals (TTIs) of various or scalable lengths (e.g., containing a variable number of OFDM symbols and / or a continuously variable absolute time).
[0078] gNBs 180a, 180b, and 180c can be configured to communicate with WTRUs 102a, 102b, and 102c in standalone and / or non-standalone configurations. In standalone configuration, WTRUs 102a, 102b, and 102c can communicate with gNBs 180a, 180b, and 180c without accessing other RANs (e.g., eNode-Bs 160a, 160b, and 160c). In standalone configuration, WTRUs 102a, 102b, and 102c can utilize one or more of gNBs 180a, 180b, and 180c as mobility anchors. In standalone configuration, WTRUs 102a, 102b, and 102c can communicate with gNBs 180a, 180b, and 180c using signals in unlicensed frequency bands. In a non-standalone configuration, WTRUs 102a, 102b, and 102c can communicate / connect with gNBs 180a, 180b, and 180c, while also communicating / connecting with another RAN such as eNode-Bs 160a, 160b, and 160c. For example, WTRUs 102a, 102b, and 102c can implement DC principles to communicate substantially simultaneously with one or more gNBs 180a, 180b, and 180c, as well as one or more eNode-Bs 160a, 160b, and 160c. In a non-standalone configuration, eNode-Bs 160a, 160b, and 160c can act as mobility anchors for WTRUs 102a, 102b, and 102c, and gNBs 180a, 180b, and 180c can provide additional coverage and / or throughput for serving WTRUs 102a, 102b, and 102c.
[0079] Each of gNBs 180a, 180b, and 180c can be associated with a specific cell (not shown) and can be configured to handle radio resource management decisions, handover decisions, user scheduling in UL and / or DL, network slicing support, dual connectivity, interoperability between NR and E-UTRA, routing user plane data to User Plane Functions (UPF) 184a and 184b, and routing control plane information to Access and Mobility Management Functions (AMF) 182a and 182b, etc. FIG. 1D As shown, gNB 180a, 180b, and 180c can communicate with each other on the Xn interface.
[0080] FIG. 1DThe CN 115 shown may include at least one AMF 182a, 182b, at least one UPF 184a, 184b, at least one Session Management Function (SMF) 183a, 183b, and possibly a Data Network (DN) 185a, 185b. Although each of the foregoing elements is depicted as part of the CN 115, it should be understood that any of these elements may be owned and / or operated by an entity other than a CN operator.
[0081] AMF 182a and 182b can connect to one or more gNBs 180a, 180b, and 180c in RAN 113 via the N2 interface and can act as control nodes. For example, AMF 182a and 182b can be responsible for authenticating users of WTRU 102a, 102b, and 102c, supporting network slicing (e.g., handling different PDU sessions with different requirements), selecting specific SMF 183a and 183b, managing registration areas, terminating NAS signaling, mobility management, and so on. AMF 182a and 182b can use network slicing to customize CN support for WTRU 102a, 102b, and 102c based on the service types used by WTRU 102a, 102b, and 102c. For example, different network slices can be established for different use cases, such as services relying on Ultra Reliable Low Latency Time (URLLC) access, services relying on Enhanced Massive Mobile Broadband (eMBB) access, services for Machine Type Communication (MTC) access, and / or so on. AMF 182a and 182b can provide control plane functions for handover between RAN 113 and other RANs (not shown) employing other radio technologies such as LTE, LTE-A, LTE-A Pro and / or non-3GPP access technologies such as WiFi.
[0082] SMFs 183a and 183b can connect to AMFs 182a and 182b in CN 115 via the N11 interface. SMFs 183a and 183b can also connect to UPFs 184a and 184b in CN 115 via the N4 interface. SMFs 183a and 183b can select and control UPFs 184a and 184b, and configure the routing of services through UPFs 184a and 184b. SMFs 183a and 183b can perform other functions, such as managing and allocating UE IP addresses, managing PDU sessions, controlling policy enforcement and QoS, and providing downlink data notifications. PDU session types can be IP-based, non-IP-based, Ethernet-based, etc.
[0083] UPF 184a and 184b can be connected to one or more gNBs 180a, 180b, and 180c in RAN 113 via the N3 interface. This N3 interface provides WTRU 102a, 102b, and 102c with access to packet-switched networks (such as Internet 110) to facilitate communication between WTRU 102a, 102b, 102c and IP-enabled devices. UPF 184 and 184b can perform other functions such as routing and forwarding packets, enforcing user plane policies, supporting multi-homed PDU sessions, handling user plane QoS, buffering downlink packets, and providing mobility anchoring.
[0084] CN 115 can facilitate communication with other networks. For example, CN 115 may include, or be able to communicate with, an IP gateway (e.g., an IP Multimedia Subsystem (IMS) server) that acts as an interface between CN 115 and PSTN 108. Furthermore, CN 115 can provide WTRUs 102a, 102b, and 102c with access to other networks 112, which may include other wired and / or wireless networks owned and / or operated by other service providers. In one embodiment, WTRUs 102a, 102b, and 102c may be connected to local data networks (DNs) 185a and 185b via the N3 interface to UPFs 184a and 184b and the N6 interface between UPFs 184a and 184b and DNs 185a and 185b.
[0085] Given FIG. 1A-1D as well as FIG. 1A-1D The functions described herein, including one or more of the following, can be performed by one or more emulation devices (not shown): WTRU 102a-d, base station 114a-b, eNode-B 160a-c, MME 162, SGW 164, PGW 166, gNB 180a-c, AMF 182a-b, UPF 184a-b, SMF183a-b, DN 185a-b, and / or one or more other devices described herein. An emulation device can be one or more devices configured to emulate one or more of the functions described herein. For example, an emulation device can be used to test other devices and / or simulate network and / or WTRU functions.
[0086] Simulation devices can be designed to perform tests on one or more other devices in laboratory and / or carrier network environments. For example, one or more simulation devices can perform one or more or all functions while being fully or partially implemented and / or deployed as part of a wired and / or wireless communication network to test other devices within the communication network. One or more simulation devices can perform one or more or all functions while being temporarily implemented / deployed as part of a wired and / or wireless communication network. Simulation devices can be directly coupled to another device for testing purposes and / or can perform tests using over-the-air wireless communication.
[0087] One or more simulation devices may perform one or more functions, including all functions, rather than being implemented / deployed as part of a wired and / or wireless communication network. For example, simulation devices may be used to test test scenarios in laboratory and / or non-deployment (e.g., testing) wired and / or wireless communication networks to implement the testing of one or more components. One or more simulation devices may be test devices. Simulation devices may transmit and / or receive data using direct RF coupling and / or wireless communication via RF circuitry (e.g., which may include one or more antennas).
[0088] This application describes various aspects, including tools, features, examples, models, methods, etc. Many of these aspects are described in detail, and often in a manner that may seem restrictive, at least to illustrate individual characteristics. However, this is for clarity and not to limit the application or scope of those aspects. In fact, all the different aspects can be combined and interchanged to provide further aspects. Furthermore, this aspect can also be combined and interchanged with aspects described in earlier applications.
[0089] The aspects described and contemplated in this application can be implemented in many different forms. FIG. 5A to FIG. 15 Some examples can be provided, but other examples are envisioned. FIG. 5A to FIG. 15 The discussion does not limit the breadth of implementation. At least one aspect typically relates to video encoding and decoding, and at least one other aspect typically relates to transmitting a generated or encoded bitstream. These and other aspects can be implemented as methods, apparatus, computer-readable storage media having instructions stored thereon for encoding or decoding video data according to any of the methods, and / or computer-readable storage media having a bitstream generated according to any of the methods stored thereon.
[0090] In this application, the terms “reconstructed” and “decoded” are used interchangeably, the terms “pixel” and “sample” are used interchangeably, and the terms “image”, “picture” and “frame” are used interchangeably.
[0091] This document describes various methods, and each method includes one or more steps or actions for implementing the method. Unless the correct operation of the method requires a specific order of steps or actions, the order and / or use of specific steps and / or actions can be modified or combined. Furthermore, terms such as "first," "second," etc., can be used in various examples to modify elements, components, steps, operations, etc., such as, for example, "first decoding" and "second decoding." Unless specifically required, the use of such terms does not imply a reordering of the modified operations. Therefore, in this example, the first decoding does not need to be performed before the second decoding and can occur, for example, before, during, or in a time period overlapping with the second decoding.
[0092] The various methods and other aspects described in this application can be used to modify, for example... FIG. 2 and FIG. 3 The illustrated video encoder 200 and decoder 300 modules are, for example, decoding modules. Furthermore, the subject matter disclosed herein can be applied to, for example, any type, format, or version of video encoding, whether described in standards or recommendations, whether pre-existing or future-developed, and any extensions to such standards and recommendations. Unless otherwise stated or technically excluded, the aspects described in this application may be used individually or in combination.
[0093] Various numerical values are used in the examples described in this application, such as degree values, number of SAO modes, number of CCSAO modes, offset values, range values, minimum values, maximum values, number of categories, number of bits, flag values, block size, number of filters, post-filter size, number of bandwidths, Golomb codes for post-filter order, dimension, number of columns, number of samples, number of candidates, order, number of positions, constant values in formulas, subsample size, filter coefficient values, power values, exponent values, etc. These and other specific values are for illustrative purposes, and the aspects described are not limited to these specific values.
[0094] FIG. 2 This is a diagram illustrating an example video encoder (e.g., a block-based hybrid video encoder). Variations of the example encoder 200 are envisioned, but for clarity, encoder 200 is described below without describing all anticipated variations.
[0095] Before being encoded, the video sequence may undergo pre-coding (201), for example, applying color transformations to the input color image (e.g., a conversion from RGB 4:4:4 to YCbCr 4:2:0), or performing remapping of the input image components to obtain a signal distribution more resilient to compression (e.g., histogram equalization using one of the color components). Metadata may be associated with pre-processing and attached to the bitstream.
[0096] In encoder 200, as described below, the image is encoded by encoder elements. The image to be encoded is partitioned (202) and processed in units, for example, coding units (CUs). Each unit is encoded using, for example, an intra-frame mode or an inter-frame mode. When a unit is encoded in intra-frame mode, it performs intra-frame prediction (260). In inter-frame mode, motion estimation (275) and compensation (270) are performed. The encoder determines (205) which of the intra-frame or inter-frame modes is used to encode the unit and indicates the intra-frame / inter-frame decision by, for example, a prediction mode flag. For example, the prediction residual is calculated by subtracting (210) the prediction block from the original image block.
[0097] The predicted residual is then transformed (225) and quantized (230). The quantized transform coefficients, along with motion vectors and other syntax elements (such as image partitioning information), are entropy encoded (245) to output a bitstream. The encoder can skip the transform and apply quantization directly to the untransformed residual signal. The encoder can bypass both the transform and quantization, i.e., directly encode the residual without applying either the transform or quantization process.
[0098] The encoder decodes the encoded blocks to provide a reference for further prediction. The quantized transform coefficients are dequantized (240) and inversely transformed (250) to decode the prediction residuals. The image blocks are reconstructed by combining (255) the decoded prediction residuals and the prediction blocks. In-loop filters (265) are applied to the reconstructed image to perform (e.g., post-processing in a predefined order) such as deblocking / SAO (Sample Adaptive Shift) / CCSAO (Cross Component Sample Adaptive Shift) / ALF (Adaptive Loop Filter) filtering to reduce coding artifacts. The filtered image is stored at the reference image buffer (280).
[0099] FIG. 3 This is a diagram illustrating an example video decoder. In the example decoder 300, as described below, the bitstream is decoded by decoder elements. The video decoder 300 typically performs operations similar to... FIG. 2 The decode traversal is the inverse of the encoding traversal described in [the document]. Encoder 200 typically also performs video decoding as part of the encoding of video data.
[0100] Specifically, the input to the decoder includes a video bitstream, which can be generated by the video encoder 200. The bitstream is first entropy decoded (330) to obtain transform coefficients, prediction modes, motion vectors, and other encoded information. Picture partitioning information indicates how the picture is partitioned. Therefore, the decoder can partition (335) the picture based on the decoded picture partitioning information. The transform coefficients are dequantized (340) and inverse transformed (350) to decode the prediction residuals. By combining (355) the decoded prediction residuals and prediction blocks, the image blocks are reconstructed. The prediction blocks can be obtained (370) from intra-frame prediction (360) or motion-compensated prediction (i.e., inter-frame prediction) (375). An in-loop filter (365) is applied to the reconstructed image. The filtered image is stored at a reference picture buffer (380). In some examples (e.g., for a given picture), the contents of the reference picture buffer 380 on the decoder 300 side can be the same as the contents of the reference picture buffer 280 on the encoder 200 side (e.g., for the same picture).
[0101] The decoded image can undergo further post-decoding processing (385), such as inverse color transformation (e.g., a conversion from YCbCr4:2:0 to RGB4:4:4) or inverse remapping of the remapping process performed in pre-encoding processing (201). Post-decoding processing can utilize metadata derived in pre-encoding processing and signaled in the bitstream. In one example, the decoded image (e.g., after applying an in-loop filter (365) and / or after post-decoding processing (385), if post-decoding processing is used) can be sent to a display device for presentation to the user.
[0102] FIG. 4 This is a diagram illustrating examples of systems in which the various aspects and examples described herein may be implemented. System 400 may be embodied as a device including the various components described below and configured to perform one or more of the aspects described in this document. Examples of such devices include, but are not limited to, various electronic devices such as personal computers, laptop computers, smartphones, tablet computers, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected home appliances, and servers. Elements of system 400 may be embodied individually or in combination in a single integrated circuit (IC), multiple ICs, and / or discrete components. For example, in at least one example, the processing and encoder / decoder elements of system 400 are distributed across multiple ICs and / or discrete components. In various examples, system 400 is communicatively coupled to one or more other systems or other electronic devices via, for example, a communication bus or through dedicated input and / or output ports. In various examples, system 400 is configured to implement one or more of the aspects described in this document.
[0103] System 400 includes at least one processor 410 configured to execute instructions loaded thereon to implement various aspects, such as those described in this document. Processor 410 may include embedded memory, input / output interfaces, and various other circuitry as known in the art. System 400 includes at least one memory 420 (e.g., a volatile memory device and / or a non-volatile memory device). System 400 includes a storage device 440 which may include non-volatile memory and / or volatile memory, including but not limited to electrically erasable programmable read-only memory (EEPROM), read-only memory (ROM), programmable read-only memory (PROM), random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), flash memory, disk drives, and / or optical disk drives. As a non-limiting example, storage device 440 may include internal storage devices, attached storage devices (including removable and non-removable storage devices), and / or network-accessible storage devices.
[0104] System 400 includes an encoder / decoder module 430 configured to, for example, process data to provide encoded or decoded video, and the encoder / decoder module 430 may include its own processor and memory. The encoder / decoder module 430 represents one or more modules that can be included in a device to perform encoding and / or decoding functions. As is known, a device may include one or both of an encoding module and a decoding module. Furthermore, as is known to those skilled in the art, the encoder / decoder module 430 may be implemented as a separate element of system 400, or it may be incorporated within processor 410 as a combination of hardware and software.
[0105] Program code to be loaded onto processor 410 or encoder / decoder 430 to execute the various aspects described in this document may be stored in storage device 440 and subsequently loaded onto memory 420 for execution by processor 410. According to various examples, one or more of processor 410, memory 420, storage device 440, and encoder / decoder module 430 may store one or more various items during the execution of the processes described in this document. Such stored items may include, but are not limited to, input video, decoded video or portions of decoded video, bitstreams, matrices, variables, and intermediate or final results from the processing of equations, formulas, operations, and operational logic.
[0106] In some examples, the memory within processor 410 and / or encoder / decoder module 430 is used to store instructions and provide working memory for processing required during encoding or decoding. However, in other examples, external memory (e.g., the processing device may be processor 410 or encoder / decoder module 430) is used for one or more of these functions. External memory may be memory 420 and / or storage device 440, such as dynamically volatile memory and / or non-volatile flash memory. In several examples, external non-volatile flash memory is used to store, for example, the operating system of a television. In at least one example, fast external dynamically volatile memory, such as RAM, is used as working memory for video encoding and decoding operations.
[0107] Inputs to the components of system 400 may be provided by various input devices, as indicated in block 445. Such input devices include, but are not limited to: (i) an RF section that receives radio frequency (RF) signals transmitted over the air, for example by a broadcasting device; (ii) a component (COMP) input terminal (or a set of COMP input terminals); (iii) a universal serial bus (USB) input terminal; and / or (iv) a high-definition multimedia interface (HDMI) input terminal. FIG. 4 Other examples not shown include composite videos.
[0108] In various examples, as is known in the art, the input device of block 445 has associated respective input processing elements. For example, the RF section may be associated with elements suitable for: (i) selecting a desired frequency (also known as selecting a signal, or limiting the signal band to a band), (ii) down-converting the selected signal, (iii) band-limiting it again to a narrower band to select, for example, a signal band that may be referred to as a channel in some examples, (iv) demodulating the down-converted and band-limited signal, (v) performing error correction, and / or (vi) demultiplexing to select a desired data packet stream. The RF section of various examples includes one or more elements to perform these functions, such as frequency selectors, signal selectors, band limiters, channel selectors, filters, downconverters, demodulators, error correctors, and demultiplexers. The RF section may include a tuner that performs various of these functions, including, for example, down-converting a received signal to a lower frequency (e.g., intermediate frequency or near-baseband frequency) or baseband. In one set-top box example, the RF section and its associated input processing elements receive RF signals transmitted via a wired (e.g., cable) medium and perform frequency selection by filtering, down-converting, and re-filtering to the desired frequency band. Various examples rearrange the order of the above (and other) components, remove some of these components, and / or add other components that perform similar or different functions. Adding components may include inserting components between existing components, such as, for example, inserting an amplifier and an analog-to-digital converter. In various examples, the RF section includes an antenna.
[0109] USB and / or HDMI terminals may include their respective interface processors for connecting system 400 to other electronic devices across USB and / or HDMI connections. It should be understood that various aspects of input processing (e.g., Reed Solomon error correction) can be implemented as needed, for example, within a separate input processing IC or within processor 410. Similarly, various aspects of USB or HDMI interface processing can be implemented as needed, either within a separate interface IC or within processor 410. Demodulated, error-corrected, and demultiplexed streams are provided to various processing elements, including, for example, processor 410 and encoder / decoder 430, to operate in conjunction with memory and storage elements to process the data streams as needed for presentation on the output device.
[0110] Various components of system 400 can be housed within an integrated housing. Within the integrated housing, various components can be interconnected and transmit data between them using suitable connection arrangements 425 (e.g., internal buses as known in the art, including inter-IC (I2C) buses, wiring, and printed circuit boards).
[0111] System 400 includes a communication interface 450 that enables communication with other devices via a communication channel 460. The communication interface 450 may include, but is not limited to, a transceiver configured to transmit and receive data via the communication channel 460. The communication interface 450 may include, but is not limited to, a modem or network interface card (NIC), and the communication channel 460 may be implemented, for example, in a wired and / or wireless medium.
[0112] In various examples, data is streamed to or otherwise provided to system 400 using a wireless network such as a Wi-Fi network (e.g., IEEE 802.11 (IEEE refers to the Institute of Electrical and Electronics Engineers)). In these examples, the Wi-Fi signal is received via a communication channel 460 and a communication interface 450 suitable for Wi-Fi communication. The communication channel 460 in these examples is typically connected to an access point or router that provides access to external networks, including the Internet, to allow streaming applications and other over-the-top communications. Other examples use a set-top box that delivers data via an HDMI connection to input block 445 to provide streaming data to system 400. Still other examples use an RF connection to input block 445 to provide streaming data to system 400. As indicated above, various examples provide data in a non-streaming manner. Furthermore, various examples use wireless networks other than Wi-Fi, such as cellular networks or Bluetooth® networks.
[0113] System 400 can provide output signals to various output devices, including display 475, speaker 485, and other peripheral devices 495. Display 475, in various examples, includes one or more of, for example, a touchscreen display, an organic light-emitting diode (OLED) display, a curved display, and / or a foldable display. Display 475 can be used in televisions, tablets, laptops, mobile phones, or other devices. Display 475 can also be integrated with other components (e.g., as in a smartphone) or separate (e.g., as in an external monitor for a laptop). In various examples, other peripheral devices 495 include one or more of a standalone digital video disc (or digital multifunction disc) (both terms refer to DVDs), a disc player, a stereo system, and / or a lighting system. Various examples use one or more peripheral devices 495 that provide functionality based on the output of system 400. For example, a disc player performs the function of playing the output of system 400.
[0114] In various examples, signaling such as AV links, Consumer Electronics Control (CEC), or other communication protocols enabling device-to-device control with or without user intervention is used to communicate control signals between system 400 and display 475, speaker 485, or other peripheral devices 495. Output devices can be communicatively coupled to system 400 via dedicated connections through their respective interfaces 470, 480, and 490. Alternatively, output devices can be connected to system 400 via communication interface 450 using communication channel 460. Display 475 and speaker 485 can be integrated into a single unit with other components of system 400 in electronic devices such as, for example, televisions. In various examples, display interface 470 includes display drivers, such as, for example, timing controller (TCon) chips.
[0115] For example, if the RF section of input 445 is part of a standalone set-top box, then display 475 and speaker 485 can alternatively be separate from one or more other components. In various examples where display 475 and speaker 485 are external components, the output signal can be provided via dedicated output connections, including, for example, an HDMI port, a USB port, or a COMP output.
[0116] The example can be implemented by computer software implemented by processor 410, or by hardware, or by a combination of hardware and software. As a non-limiting example, the example can be implemented by one or more integrated circuits. As a non-limiting example, memory 420 can be of any type suitable for the technical environment and can be implemented using any suitable data storage technology, such as optical storage devices, magnetic storage devices, semiconductor-based memory devices, fixed memory, and removable memory. As a non-limiting example, processor 410 can be of any type suitable for the technical environment and can encompass one or more of microprocessors, general-purpose computers, special-purpose computers, and processors based on multi-core architectures.
[0117] Various implementations involve decoding. As used in this application, "decoding" can include, for example, all or part of a process performed on a received encoded sequence to produce a final output suitable for display. In various examples, such a process includes one or more processes typically performed by a decoder, such as entropy decoding, inverse quantization, inverse transform, and differential decoding. In various examples, such a process may also, or alternatively, include processes performed by decoders of various implementations described in this application, such as determining a post-filter order associated with a region (e.g., a slice) or a sub-region of an image (e.g., a coding tree unit (CTU)), wherein the post-filter order indicates a post-processing order; receiving the post-filter order in a slice header or image header; determining a luma post-filter order or a chroma post-filter order; determining a luma post-filter order and applying that luma post-filter order to the chroma post-filter order; generating coding costs and distortion; determining the rate distortion cost of each CTU of a slice when configured to receive the post-filter order in a slice header; determining the rate distortion cost of each CTU of an image when configured to receive the post-filter order in an image header; and determining the rate distortion cost based on the distortion of the current block with its coding parameters, the associated rate or cost, and / or parameters derived from quantization parameters, etc.
[0118] As further examples, in one example, "decoding" refers only to entropy decoding; in another example, "decoding" refers only to differential decoding; and in yet another example, "decoding" refers to a combination of entropy decoding and differential decoding. It will be clear, and is considered well understood by one of ordinary skill in the art, whether the phrase "decoding process" is intended to refer specifically to a subset of operations or to a broader decoding process, based on the specific context of the description.
[0119] Various implementations involve encoding. In a manner similar to the discussion above regarding “decoding,” “encoding,” as used in this application, can include, for example, all or part of a process performed on an input video sequence to produce an encoded bitstream. In various examples, such a process includes one or more processes typically performed by an encoder, such as partitioning, differential coding, transform, quantization, and entropy coding. In various examples, such a process also or alternatively includes processes performed by encoders of various implementations described in this application, such as determining a post-filter order associated with a region (e.g., a slice) or a sub-region of a picture (e.g., a coding tree unit (CTU)), wherein the post-filter order indicates a post-processing order; sending the post-filter order in a region header (e.g., a slice header) or picture header; determining a luma post-filter order or a chroma post-filter order; determining a luma post-filter order and applying that luma post-filter order to the chroma post-filter order; generating coding costs and distortion; determining the rate distortion cost of each CTU of a slice when configured to send the post-filter order in a slice header; determining the rate distortion cost of each CTU of a picture when configured to send the post-filter order in a picture header; and determining the rate distortion cost based on the distortion of the current block with the current block coding parameters, the associated rate or cost, and / or parameters derived from the quantization parameters.
[0120] As further examples, in one example, "encoding" refers only to entropy encoding; in another, "encoding" refers only to differential encoding; and in yet another, "encoding" refers to a combination of differential and entropy encoding. It will be clear, and is considered well understood by those skilled in the art, whether the phrase "encoding process" is intended to specifically refer to a subset of operations or to refer generally to the broader decoding process, based on the specific context of the description.
[0121] Note that the syntax elements used in this article (such as ph_postproc_order, rdo_postprocessing_order_picture(), sps_sao_enabled_flag, pps_sao_info_in_ph_flag, ph_sao_luma_enabled_flag, sps_ccsao_enabled_flag, ph_cc_sao_y_enabled_flag, pps_dbf_info_in_ph_flag, ph_deblocking_params_present_flag, sps_alf_enabled_flag, pps_alf_info_in_ph_flag, ph_alf_enabled_flag, etc., which may be detailed in the example logic, algorithms, equations, etc.) are descriptive terms. Therefore, their use does not preclude the use of other syntax element names.
[0122] When a diagram is presented as a flowchart, it should be understood that it also provides a block diagram of the corresponding device. Similarly, when a diagram is presented as a block diagram, it should be understood that it also provides a flowchart of the corresponding method / process.
[0123] The implementations and aspects described herein can be implemented, for example, in a method or process, apparatus, software program, data stream, or signal. Even if discussed only in the context of a single implementation (e.g., discussed only as a method), the implementation of the discussed features can also be implemented in other forms (e.g., apparatus or program). An apparatus can be implemented, for example, in appropriate hardware, software, and firmware. A method can be implemented, for example, in a processor, which generally refers to a processing device, including, for example, a computer, microprocessor, integrated circuit, or programmable logic device. A processor also includes communication devices, such as, for example, a computer, a cellular phone, a portable / personal digital assistant (“PDA”), and other devices that facilitate information communication between end users.
[0124] References to “an example” or “an example” or “an implementation” or “an implementation”, and their other variations, mean that the specific feature, structure, characteristic, etc., described in connection with the example is included in at least one example. Therefore, the appearance of the phrase “in an example” or “in the example” or “in an implementation” or “in the implementation”, and any other variations appearing in different places throughout the application, do not necessarily refer to the same example.
[0125] Additionally, this application may refer to "determining" various information pieces. Determining information may include, for example, one or more of estimation information, calculation information, prediction information, or information retrieved from memory. Obtaining may include receiving, retrieving, constructing, generating, and / or determining.
[0126] Furthermore, this application may refer to "accessing" various information fragments. Accessing information may include, for example, receiving information, retrieving information (e.g., from memory), storing information, moving information, copying information, calculating information, determining information, predicting information, or estimating information, or one or more of these.
[0127] Additionally, this application may refer to "receiving" various pieces of information. Like "access," "receiving" is intended to be a broad term. Receiving information may include, for example, accessing information or retrieving information (e.g., from memory) or one or more of it. Furthermore, "receiving" is generally referred to in one way or another during operations such as, for example, storing information, processing information, transmitting information, moving information, copying information, erasing information, calculating information, determining information, predicting information, or estimating information.
[0128] It will be understood that the use of any of the following “ / ”, “and / or”, and “at least one”, for example, in the cases of “A / B”, “A and / or B”, and “at least one of A and B”, is intended to include selecting only the first listed option (A), or only the second listed option (B), or selecting both options (A and B). As another example, in the cases of “A, B, and / or C” and “at least one of A, B, and C”, such wording is intended to include selecting only the first listed option (A), or only the second listed option (B), or only the third listed option (C), or only the first and second listed options (A and B), or only the first and third listed options (A and C), or only the second and third listed options (B and C), or selecting all three options (A, B, and C). As will be apparent to those skilled in the art and related fields, this can be extended to as many items as listed.
[0129] Furthermore, as used in this article, the word “signal” refers, among other things, specifically to instructing something to the corresponding decoder. Encoder signals can include, for example, current parameters (such as offsets, flags, and indices) and filter parameters (including post-filter ordering (e.g., rdo_postprocessing_order_slice, rdo_postprocessing_order_picture), which can be signaled at various levels (e.g., picture header, region or slice header, sub-region or CTU, APS), etc. In this way, in the example, the same parameters can be used on both the encoder and decoder sides. Thus, for example, the encoder can transmit (explicitly signal) specific parameters to the decoder so that the decoder can use the same specific parameters. Conversely, if the decoder already has specific parameters as well as other parameters, signaling can be used without transmission (implicitly signal) to simply allow the decoder to know and select specific parameters. Bit savings are achieved in various examples by avoiding the transmission of any actual functionality. It will be appreciated that signaling can be implemented in many ways. For example, in various examples, one or more syntax elements, flags, etc., are used to signal information to the corresponding decoder. Although the verb form of the word "signal" was mentioned above, the word "signal" can also be used as a noun in this text.
[0130] As will be apparent to those skilled in the art, the implementation can generate various signals, which are formatted to carry information, for example, that can be stored or transmitted. This information may include, for example, instructions for performing a method, or data generated by one of the implementations. For example, the signal may be formatted to carry a bit stream of the example. Such a signal may be formatted as, for example, electromagnetic waves (e.g., using the radio frequency portion of the spectrum) or baseband signals. Formatting may include, for example, encoding the data stream and modulating a carrier wave with the encoded data stream. The information carried by the signal may be, for example, analog or digital information. As is known, the signal can be transmitted via a variety of different wired or wireless links. The signal may be stored on, accessed from, or received from a processor-readable medium.
[0131] This document describes numerous examples. Features of the examples may be provided individually or in any combination across various claim classes and types. Furthermore, examples may include one or more of the features, devices, or aspects described herein, individually or in any combination across various claim classes and types. For example, the features described herein may be implemented in a bitstream or signal including information generated as described herein. According to any of the described embodiments, this information may allow a decoder to decode the bitstream, encoder, bitstream, and / or decoder. For example, the features described herein may be implemented by creating and / or transmitting and / or receiving and / or decoding a bitstream or signal. For example, the features described herein may be implemented as a method, process, apparatus, medium storing instructions, medium storing data, or signal. For example, the features described herein may be implemented by a TV, set-top box, mobile phone, tablet computer, or other electronic device performing decoding. The TV, set-top box, mobile phone, tablet computer, or other electronic device may display (e.g., using a monitor, screen, or other type of display) the resulting image (e.g., an image reconstructed from the residual of a video bitstream). The TV, set-top box, mobile phone, tablet computer, or other electronic device may receive a signal including an encoded image and perform decoding.
[0132] As described in this article, storing the post-processing order in a dedicated code within the slice / image header can involve, for example, corresponding to... FIG. 2 265 and on the encoder side FIG. 3 The 365-degree in-loop filter on the decoder side. The post-processing sequence stored in the header of the region (e.g., a slice) and / or image, using dedicated code, can involve deblocking filtering, Cross Component Sample Adaptive Offset (CCSAO), for example, as an extension of Sample Adaptive Offset (SAO) and / or Adaptive Loop Filtering (ALF). The terms region and slice are used interchangeably.
[0133] Block-based intra / inter-frame prediction and transform coding, such as in conjunction with quantization, can introduce various artifacts at medium and / or low bit rates, such as block artifacts, ringing artifacts, and / or color aberrations. Video coding such as HEVC and VVC can implement in-loop filters (e.g., Sample Adaptive Offset (SAO) and Cross Component Sample Adaptive Offset (CCSAO)) to, for example, reduce artifacts. For example, history-based CCSAO can be implemented (e.g., in Enhanced Coding Models (ECM)) to benefit from parameter sets that may be effective in previously encoded / decoded slices. Current parameters can be encoded and signaled in a header such as a picture header, region or slice header, sub-region or coding tree level (CTU), or adaptive parameter set (APS). The terms sub-region and CTU are used interchangeably.
[0134] An optimized order of post-filters may or may not exist. In some examples, the order of post-filters can be fixed. The process / algorithm may or may not search for and / or find a better order that reduces performance rate distortion.
[0135] As described herein, storing the post-processing order in a dedicated code within the header of a slice / image can provide an indication (e.g., signaling) of the order in which post-filters are applied. This indication can specify the order at the encoder and decoder used for post-processing (e.g., the correct order).
[0136] Sample Adaptive Offset (SAO) can be implemented, for example, in HEVC. SAO can be used to reduce the average sample distortion of a region, for example, by classifying the region samples into multiple classes using a selected classifier (e.g., first), obtaining an offset for each class, and then adding that offset to each sample of that class. The classifier index and / or the region offset can be encoded into a bitstream. In SAO, the offset and / or parameters can be signaled at the CTU level.
[0137] SAO can be implemented in, for example, HEVC and VVCA. Similarly, SAO can be implemented in VVC and HEVC.
[0138] A coding tree unit (CTU) (e.g., when enabled) can be encoded using multiple (e.g., three (3)) SAO modes (SaoTypeIdx), such as inactive (OFF), edge offset (EO), and / or band offset (BO). For example, in the case of EO or BO, a parameter set (e.g., a parameter set) can be encoded for each channel (Y, U, V). This parameter set can be shared with neighboring CTUs (e.g., SAO MERGE flags). The SAO mode can be the same for Cb and Cr components.
[0139] FIG. 5A-5C The illustration shows an example of reconstructing sample class determination in the case of EO mode. For example, in the case of EO, each reconstructed sample can be classified into multiple (e.g., NC=5) classes (e.g., sao_eo_class). The classification of (one or more) classes can depend on the local gradient following the direction (e.g., one) signaled by each CTU (e.g., EO_0, EO_90, EO_135, or EO_45 corresponding to the direction of 0 degrees, 90 degrees, 135 degrees, or 45 degrees), such as FIG. 5A-5C As illustrated in the example, multiple (e.g., NC-1) offset values can be encoded, for example, encoding once for each category, where a category can have an offset equal to zero. The offset sign can be derived from the category (e.g., implicitly derived).
[0140] FIG. 6 This describes a pixel range that can be uniformly divided into multiple (e.g., 32) frequency bands, for example, in the case of BO mode. For example, the uniformly divided pixel range can be from 0 to 255 (e.g., in 8 bits). In one example, for example in the case of BO, the sample range of the value of component "c" (e.g., 0…255, in 8 bits) can be divided (e.g., uniformly divided) into multiple (e.g., N(c) = 32) frequency bands. Sample values belonging to multiple (e.g., (NC-1) = 4) consecutive frequency bands can be modified, for example, by adding an offset off(n). FIG. 6 An example of four (4) consecutive frequency bands is shown. The offset values of the band symbols (e.g., the offset values of (NC-1) band symbols) and / or the starting frequency band can be encoded, for example, each of the (NC-1) frequency bands can be encoded once, where the remaining frequency bands can have an offset equal to zero.
[0141] For example, a lookup table (LUT) 'lut' can be built for each component 'c', including the offset for each frequency band. SAO [c]'. For example, in addition to NC-1 consecutive values, the LUT can include zeros, which include, for example, offsets added to the reconstructed sample according to equation (1): Referring to equation (1), the parameter rec[x] can be the reconstructed sample of the current component (c) at position 'x' (e.g., before applying SAO). The parameter s can be equal to (BD(c) – 5), where BD(c) can be the bit depth of component 'c'. The parameter lut SAO [c] can include the offset of component 'c'.
[0142] For example, in the case of EO or BO, the SAO pattern and / or offset may not be encoded. The SAO pattern and / or offset can be copied from the adjacent upper or left CTU (e.g., in merged mode).
[0143] For example, as described herein, SAO flags can be transmitted and / or signaled. For example, the Picture Parameter Set (PPS) can be transmitted and / or signaled according to the following logic: , For example, the Sequence Parameter Set (SPS) can be transmitted and / or signaled according to the following logic: , For example, the image header can be transmitted and / or signaled according to the following logic: , For example, the slice header can be transmitted and / or signaled according to the following logic: .
[0144] The flag pps_no_pic_partition_flag (e.g., equal to 1) can specify that image partitioning is not applied to images referencing PPS (e.g., each one). The flag pps_no_pic_partition_flag (e.g., equal to 0) can specify that images referencing PPS (e.g., each one) may be partitioned into more than one tile or slice.
[0145] The flag `pps_sao_info_in_ph_flag` (e.g., equal to 1) specifies that SAO filter information may exist in the PH syntax structure and / or not exist in the slice header referencing a PPS that does not contain a PH syntax structure. The flag `pps_sao_info_in_ph_flag` (e.g., equal to 0) specifies that SAO filter information does not exist in the PH syntax structure and / or may exist in the slice header referencing a PPS. For example, when the flag `pps_sao_info_in_ph_flag` is not present, its value can be inferred (e.g., it should be equal to 0).
[0146] A Picture Parameter Set (PPS) can be a syntax structure that may include syntax elements applied to zero or more complete encoded pictures, such as those determined by syntax elements found in (e.g., each) the picture header.
[0147] The flag `sps_sao_enabled_flag` (e.g., equal to 1) can specify that SAO is enabled for coding layer video sequences. The flag `sps_sao_enabled_flag` (e.g., equal to 0) can specify that SAO is disabled for coding layer video sequences.
[0148] A sequence parameter set (SPS) can be a grammatical structure that includes grammatical elements applied to zero or more complete coding layer video sequences, such as those defined by the content of grammatical elements found in the PPS, which is referenced by grammatical elements found in (e.g., each) picture header.
[0149] The flag `ph_sao_luma_enabled_flag` (e.g., equal to 1) can specify that SAO is enabled for the luma component of the current image. The flag `ph_sao_luma_enabled_flag` (e.g., equal to 0) can specify that SAO is disabled for the luma component of the current image. For example, if the flag `ph_sao_luma_enabled_flag` does not exist, it can be inferred that the flag `ph_sao_luma_enabled_flag` is enabled (e.g., it should be equal to 0).
[0150] The flag `ph_sao_chroma_enabled_flag` (e.g., equal to 1) can specify that SAO is enabled for the chroma components of the current image. The flag `ph_sao_chroma_enabled_flag` (e.g., equal to 0) can specify that SAO is disabled for the chroma components of the current image. For example, if the flag `ph_sao_chroma_enabled_flag` does not exist, it can be inferred that the flag `ph_sao_chroma_enabled_flag` is equal to 0.
[0151] The `slice_sao_luma_flag` (e.g., equal to 1) can specify that SAO is used for the luma component in the current slice. The `slice_sao_luma_flag` (e.g., equal to 0) can specify that SAO is not used for the luma component in the current slice.
[0152] The `slice_sao_chroma_flag` (e.g., equal to 1) can specify that SAO is used for the chroma components in the current slice. The `slice_sao_chroma_flag` (e.g., equal to 0) can specify that SAO is not used for the chroma components in the current slice.
[0153] FIG. 7This is a diagram illustrating an example of a decoding workflow (e.g., a modified SAO process when CCSAO is applied). Cross Component Sample Adaptive Offset (CCSAO) can be implemented, for example, in ECM. CCSAO can be used to refine reconstructed chroma samples. CCSAO can (e.g., similar to SAO) classify reconstructed samples into different categories. CCSAO can derive (e.g., one) offset for (e.g., each) category. CCSAO can add the offset to the reconstructed samples in the categories for which CCSAO derives the offset. CCSAO can differ from SAO. SAO can (e.g., only) use one (e.g., a single) luma / chroma component of the current sample as input. CCSAO can utilize multiple (e.g., all three) components to classify the current sample into different categories. Output samples from the deblocking filter can be used as input to CCSAO, for example, to facilitate parallel processing.
[0154] CCSAO can (e.g., only) use BO to enhance the quality of reconstructed samples. For a given luminance / chrominance sample, multiple (e.g., three) candidate samples can be selected to classify the given sample into different categories, such as (e.g., one) juxtaposed Y samples, (e.g., one) juxtaposed U samples, and (e.g., one) juxtaposed V samples. The sample values of the (e.g., three) selected samples can be classified into (e.g., three) different frequency bands, for example, Joint Index i The category of a given sample can be represented. For a reconstructed sample that falls into that category, an offset can be signaled (e.g., a) and added to the reconstructed sample that falls into that category, which can be formulated / implemented according to a set of equations (e.g., as shown in equations (2) – (6)). .
[0155] Refer to equations (2) – (6). It can be three selected, juxtaposed samples used to classify the current sample; They can be applied to respectively The number of equally divided frequency bands across the entire range; BD It can be the internal encoding bit depth; and It can be reconstructed samples before and after applying CCSAO; and This can be the value of the CCSAO offset applied to the i-th BO category. Luminance samples can be juxtaposed from multiple (e.g., nine (9)) candidate locations. The positions of the juxtaposed chrominance samples can be fixed, for example, as determined by... FIG. 8 The example depicted in the text.
[0156] FIG. 8The illustration shows an example of candidate locations for a CCSAO classifier. For example, CCSAO can apply different classifiers to various local regions to further enhance (e.g., overall) image quality. The parameters (e.g., location) of the classifier can be signaled (e.g., for each) at the image level. , , , (and offset). The classifier to be used can be signaled and / or switched at the CTB level (e.g., explicitly). Settings can be set for (e.g., each) classifier. The maximum value can be set, for example, to {16, 4, 4}, and / or the offset can be limited to a range such as [-15, 15]. In some examples, up to four (4) classifiers can be used per frame.
[0157] FIG. 9 This diagram illustrates an example of joint clipping after adding SAO / BIF / CCSAO offsets to the input samples. In some examples, the SAO, BIF, and / or CCSAO offsets can be computed in parallel, added to the reconstructed chroma samples, and then jointly clipped, as shown below. FIG. 9 As depicted in [the text]. The edge-based classifier of CCSAO can (e.g., similar to the edge classifier of SAO in VVC) classify samples using four one-dimensional (1-D) orientation patterns, such as horizontal, vertical, 135° diagonal, and 45° diagonal (e.g., as in [the text]). FIG. 10 (As illustrated in the example).
[0158] FIG. 10 The illustration shows examples of four 1-D orientation patterns used for CCSAO EO sample classification: horizontal (EO class = 0), vertical (EO class = 1), 135° diagonal, and 45° diagonal. FIG. 10 As shown in the example, samples can be classified for (e.g., each) 1-D pattern based on, for example, the sample difference between a luminance sample value labeled “c” and two neighboring luminance samples labeled “a” and “b” along the selected 1-D pattern.
[0159] CCSAO can (e.g., similar to SAO) involve the encoder using rate-distortion optimization (RDO) to determine (e.g., the optimal) 1-D orientation pattern. The encoder can signal additional information in (e.g., each) classifier / ensemble. The sample differences “ac” and “bc” can be compared with (e.g., a predefined) threshold (Th) to derive (e.g., the final) “class_idx” information.
[0160] The encoder can, for example, select (e.g., optimal) "Th" values from a predefined threshold array based on RDO. The index of the "Th" array can be signaled. Differences (e.g., additional differences) may exist between a CCSAO-based edge classifier and a SAO edge classifier, such as in VVC. For example, chroma samples in a CCSAO can have edge information derived using co-localized luminance samples (e.g., samples "a", "c", and "b" can be co-localized luminance samples). Chroma samples in a SAO can have edge information derived using their own neighboring samples.
[0161] The edge-based classifier process can be formulated / implemented according to a set of equations (e.g., as shown in equations (7) – (10)): , Referring to equation (3C), the variable “ can be derived, for example, from equation (11). i B ": , The current sample "cur" can be the currently being processed sample. Parameter and These can be samples from collaborative localization. When brightness samples are processed, the parameters... and They can be cooperative positioning. Samples and Sample. When the chromaticity ( C b When the sample is processed, the parameters and They can be cooperative positioning. Y Samples and Sample. When the chromaticity ( C r When the sample is processed, the parameters and They can be cooperative positioning. Y Samples and sample.
[0162] Encoder signals (e.g., sample "cur") "", One of them () can be used, for example, to derive frequency band information based on RDO.
[0163] CCSAO offsets and / or parameters can be transmitted. For example, as disclosed herein, parameters and / or flags can be transmitted and / or signaled.
[0164] Flags such as sps_ccsao_enabled_flag can be used to send and / or signal to SPS.
[0165] The image header can be transmitted and / or signaled according to the following: If the `sps_ccsao_enabled_flag` enables the flag and the "SAO information in the image header" flag: .
[0166] The slice header can be transmitted and / or signaled according to the following: If sps_ccsao_enabled_flag is enabled For example, for each of the components Y, Cr, Cb: .
[0167] The parameter `sps_ccsao_enabled_flag` can be a flag specifying that SPS CCSAO is enabled (e.g., equal to 1). The parameter `ph_cc_sao_y_enabled_flag` can be a flag specifying that the image header CCSAO is enabled in Luminance (e.g., equal to 1). The parameter `ph_cc_sao_cr_enabled_flag` can be a flag specifying that the image header CCSAO is enabled in Cr (e.g., equal to 1). The parameter `ph_cc_sao_cb_enabled_flag` can be a flag specifying that the image header CCSAO is enabled in Cb (e.g., equal to 1). The parameter `slice_ccsao_y_enabled_flag` can be a flag specifying that slice CCSAO is enabled in Y (e.g., equal to 1). The parameter `slice_ccsao_cr_enabled_flag` can be a flag specifying that slice CCSAO is enabled in Cr (e.g., equal to 1). The parameter `slice_ccsao_cb_enabled_flag` can be a flag specifying whether slice CCSAO is enabled in Cb (e.g., equal to 1). The parameter `ccsao_set_num` can be the number of sets enabled. In some examples, it can indicate the maximum number of sets (e.g., `MAX_CCSAO_SET_NUM=4`). The parameter `setType` can be a flag specifying whether the offset type is edge (e.g., equal to 1) or band (e.g., equal to 0). FIG. 8As shown in the diagram, a value can be assigned to the parameter ccsao_cand_pos, which specifies (e.g., the best) candidate position y among nine (9) positions. The parameter ccsao_band_num_y can specify the band luminance. The parameter ccsao_band_num_c can specify the band chrominance (cb). The parameter ccsao_band_num_u can specify the band Cb. The parameter ccsao_band_num_v can specify the band Cr. The parameter ccsao_offset_abs can specify the absolute value of the offset. The parameter ccsao_offset_sign can specify the sign of the offset value.
[0168] History-based CCSAO can be implemented. Strong temporal correlations can exist between CCSAO offsets and classifier parameters across different frames. The CCSAO offsets / parameters of the current frame can be stored and / or referenced for CCSAO processes in future frames. The index can be signaled in the slice header to indicate which stored offset / parameter set was inherited from previous slices in the current slice.
[0169] For example, in VVC, an adaptive loop filter (ALF) with block-based filter adaptation can be applied. A filter can be selected for the luminance component. For example, one of 25 filters can be selected for each (e.g., each) 4×4 block. Alternatively, the filter can be selected based on the direction and activity of local gradients.
[0170] FIG. 11 The illustration shows an example of an ALF filter shape (e.g., chroma: 5×5 rhombus, luma: 7×7 rhombus). Multiple (e.g., two) rhombus filter shapes (e.g., as shown in the image) FIG. 11 (As shown) can be used in ALF. For example... FIG. 11 As shown, a 7×7 rhombus shape can be applied to the luminance component, and / or a 5×5 rhombus shape can be applied to the chrominance component.
[0171] Block classification can be performed. For example (for the luminance component), each 4×4 block can be classified into one of 25 classes. The classification index C can be based, for example, on its orientation. D Quantitative values of activities Based on equation (12): , Horizontal, vertical and two diagonal directions D and activities The gradient can be calculated using a 1-D Laplace, for example, according to equations (13)–(16): , indexi and j It can refer to the coordinates of the top left sample within a 4×4 block. Can indicate coordinates Reconstructed samples at the location.
[0172] FIG. 12A-12D The illustration shows an example of a subsampling location used for subsampling Laplace computation. For instance, applying subsampling 1-D Laplace computation can reduce the complexity of block classification. FIG. 12A-12D As illustrated in the diagram, the same subsampling location can be used for gradient calculation in directions (e.g., all directions).
[0173] FIG. 12A The illustration shows an example of the subsampling location for calculating the vertical gradient using subsampling Laplacian. FIG. 12B The illustration shows an example of the subsampling location for calculating the horizontal gradient using subsampling Laplacian. FIG. 12C The illustration shows an example of the subsampling location for the diagonal gradient used in the subsampling Laplacian calculation. FIG. 12D The illustration shows an example of the subsampling location for the diagonal gradient used in the subsampling Laplacian calculation.
[0174] For example, the gradients in the horizontal and vertical directions can be set according to equation (17). D Maximum and minimum values: , For example, the maximum and minimum values of the gradients in the two diagonal directions can be set according to equation (18): , For example, this value can be compared with two thresholds by, for instance, according to the following logic. and To derive direction through comparison D Value: if and If both are true, then D It was set to 0. if If yes, continue from step 3; otherwise, continue from step 4. if ,but D Set to 2; otherwise D It was set to 1. if ,but D Set to 4; otherwise D It was set to 3. For example, the activity value can be calculated according to equation (19). A : , A It can be quantized (e.g., quantized to the range of 0 to 4, inclusive). Quantized values can be represented as... For example, classification may not be applied to chromaticity components.
[0175] Geometric transformations of the filter coefficients and limiting values can be performed. For example, (e.g., before filtering each 4×4 luminance block), geometric transformations such as rotation or diagonal and vertical flipping can be applied to the filter coefficients, for example, based on the gradient values calculated for the blocks. and / or the corresponding filter limiting value Performing a geometric transformation on the filter coefficients and / or limiting values is equivalent to applying the transformation to samples within the filter's support region. Performing a geometric transformation on the filter coefficients and / or limiting values can make different blocks to which ALF has been applied more similar by aligning their orientations.
[0176] Multiple (e.g., three) geometric transformations (such as diagonal, vertical flip, and rotation) can be achieved, for example, according to equations (20)–(22): diagonal: Vertical Flip: Rotation: , K can be the size of the filter, and , These are coefficient coordinates, for example, making the position... In the top left corner, and in position In the bottom right corner. For example, based on the gradient values calculated for the block, the transformation can be applied to the filter coefficients. and / or limiting value The relationship between the four gradients and transformations in the four directions can be summarized, for example, according to Table 1.
[0177] Table 1 – Examples of gradients computed by mapping to blocks and transformations Gradient value Transform g d2 g d1 and g h g v ]]> No transform g d2 <g d1 and g v <g h ]]> Diagonal g d1 g d2 and g h g v ]]> Vertical flip g d1 g d2 and g v g h ]]> Rotation
[0178] Filtering can be performed through a filtering process. It can be done on the decoder side (e.g., when ALF is enabled for CTB) on each sample within the CU. Filtering is performed to generate sample values. It can be determined according to equation (23): , Refer to equation (23). It can represent the filter coefficients for decoding. It can be used for amplitude limiting, and This can represent the decoding limiting parameters. Variables k and l can be used... and The changes between them, among which L It can represent the filter length. Limiting function. It can correspond to the function Amplification can introduce a non-linear function, which can make ALF more efficient, for example, by reducing the influence of neighboring sample values that differ too much from the current sample value.
[0179] FIG. 13A This is a system-level diagram illustrating an example of the CC-ALF process related to the SAO, Luminance ALF, and Chroma ALF processes. The Cross-Component Adaptive Loop Filter (CC-ALF) can refine (e.g., each) the chroma components using luminance sample values, for example, by applying an adaptive linear filter to the luminance channel. The output of the filtering operation can be used for chroma refinement.
[0180] FIG. 13B This is a diagram illustrating an example of a diamond-shaped filter. Filtering in CC-ALF can be achieved, for example, by using a linear diamond-shaped filter (e.g., such as...). FIG. 13B The example shown in the image is applied to the luminance channel to accomplish this. (For example, a) filter can be applied to (for example, each) chroma channel. This operation can be expressed, for example, according to equation (24): , Refer to equation (24). It could be the position of the refined chromaticity component i. It can be based on The brightness position, S i It can be the filter support region in the luminance component, and It can represent filter coefficients.
[0181] like FIG. 13A and FIG. 13B As shown, the luminance filter support can be a region juxtaposed with the current chroma sample, for example, after taking into account the spatial scaling factor between the luminance plane and the chroma plane. FIG. 13A The illustration shows an example of the placement of the CC-ALF relative to other loop filters. FIG. 13B An example of a diamond-shaped filter is illustrated.
[0182] For example, CC-ALF filter coefficients can be computed (e.g., in VVC reference software) by minimizing the mean squared error (MSE) of each chroma channel relative to the original chroma content. This can be achieved using the VVC Test Model (VTM) algorithm, which employs a coefficient derivation process similar to that used for chroma ALF. For example, the correlation matrix can be derived, and the coefficients can be computed using a Joleski decomposition solver to attempt to minimize the MSE metric. In some examples, up to eight (8) CC-ALF filters can be designed and delivered per image. The filters can be based on the CTU, for example, the results obtained for each of the two chroma channels.
[0183] A CC-ALF implementation may include one or more of the following (e.g., additional) features: the CC-ALF design can use a 3×4 diamond shape with 8 taps; seven filter coefficients can be transmitted in the APS; (e.g., each) the transmitted coefficients can have a 6-bit dynamic range and / or can be limited to values that are powers of 2; an eighth filter coefficient can be derived at the decoder (e.g., such that the sum of the filter coefficients equals 0); the APS can be referenced in the slice header; for (e.g., each) chroma components, CC-ALF filter selection can be controlled at the CTU level; and / or the boundary padding of the horizontal virtual boundary can use the same memory access mode as the luma ALF.
[0184] The reference encoder can be configured to enable (e.g., subjective) tuning via a configuration file. The VTM can (e.g., when enabled) attenuate the application of CC-ALF in regions encoded with high QP and / or either close to mid-grayscale or containing (e.g., a large amount) of high-frequency luminance, which can be done (e.g., algorithmically) by disabling the application of CC-ALF in the CTU if one or more of the following conditions are true: the slice QP value minus 1 is less than or equal to the base QP value; the number of chroma samples with local contrast greater than (1 << (bit depth – 2)) – 1 exceeds the CTU height, where local contrast can be the difference between the maximum and minimum luminance sample values within the filter-supported region; and / or more than a certain number or percentage (e.g., a quarter) of chroma samples are in the range between (1 << (bit depth – 1)) – 16 and (1 << (bit depth – 1)) + 16.
[0185] This functionality can provide some (e.g., probabilistic) guarantee that CC-ALF will not amplify artifacts introduced earlier in the decoding path (e.g., at least in part due to the lack of explicit optimization for subjective chroma quality in VTM). Some encoder implementations may not use one or more of the functionalities described herein, or may combine them with alternative strategies suitable for their coding characteristics.
[0186] Filter parameters can be signaled. ALF filter parameters can be signaled in the Adaptive Parameter Set (APS). In some examples, the APS may include up to 25 sets of luminance filter coefficients and limiting indexes, and up to eight (8) sets of chrominance filter coefficients and limiting indexes. For example, filter coefficients from different categories of the luminance component can be combined to reduce bit overhead. The index of the APS used for the current slice can be signaled, for example, in the slice header.
[0187] The limiting value can be determined from the limiting value index decoded by the AP, which allows for the use of a table of limiting values for, for example, luminance and / or chrominance components (e.g., both). The limiting value can depend on the internal bit depth. For example, the limiting value can be obtained according to equation (25): , Referring to equation (25), B can be equal to the internal bit depth. It can be a value (e.g., a predefined constant) (e.g., equal to 2.35), and N can be equal to the number of allowed limit values (e.g., 4 in VVC). AlfClip can be rounded to the nearest value, e.g., a value in the format of a power of two (2).
[0188] Multiple (e.g., up to seven (7)) APS indices can be signaled in the slice header to specify the luma filter bank for the current slice. The filtering process can be (e.g., further) controlled at the CTB level. A signaled flag can be (e.g., always) used to indicate whether an ALF is applied to the luma CTB. The luma CTB can select a filter bank from multiple (e.g., 16) fixed filter banks and / or filter banks from the APS. The filter bank index can be signaled for the luma CTB to indicate which filter bank is applied. Multiple (e.g., 16) fixed filter banks can be predefined and / or hard-coded in the encoder and / or decoder.
[0189] The APS index can be signaled in the slice header of the chroma component to indicate the chroma filter bank used for the current slice. For example, if there is more than one chroma filter bank in the APS, the filter index can be signaled for each chroma CTB at the CTB level.
[0190] For example, filter coefficients can be quantized using a norm equal to 128. This can reduce multiplication complexity. For example, bitstream consistency can be applied to ensure that coefficient values at non-center positions are within a certain range (e.g., within 2). 7 to 2 7 The range is -1, including 2. 7 and 2 7-1). The center position coefficient may not be signaled in the bitstream. The center position coefficient may (for example, be considered) equal to 128.
[0191] For example, flags and parameters can be transmitted according to the following logic. For example, a Picture Parameter Set (PPS) can be transmitted according to the following logic: , For example, the Sequence Parameter Set (SPS) can be transmitted according to the following logic: For example, the image header can be transmitted according to the following logic: , For example, the slice header can be transmitted according to the following logic: ; For example, APS can be transmitted according to the following logic: .
[0192] Deblocking filtering (DBF) can be applied to CU boundaries, transformed subblock boundaries, and / or predicted subblock boundaries. Predicted subblock boundaries can include prediction unit boundaries introduced by SbTMVP and affine modes. Transformed subblock boundaries can include transform unit boundaries introduced by SBT and ISP modes, and / or transforms attributed to (e.g., implicit) segmentation of large CUs. In HEVC, the processing order of deblocking filters can be defined as horizontal filtering for vertical edges of the image (e.g., first the entire image) and (e.g., subsequently) vertical filtering for horizontal edges. This processing order allows multiple horizontal or vertical filtering processes to be applied in parallel threads, or can be implemented on a CTB-by-CTB basis (e.g., with small processing latency).
[0193] Deblocking (e.g., deblocking in VVC) can implement one or more of the following: Deblocking can implement a deblocking filter with a filter strength that depends on the average luminance level of the reconstructed samples. Deblocking can implement deblocking tC table expansion and / or adaptation for 10-bit video. Deblocking can implement 4×4 grid deblocking for luminance. VVC deblocking can implement a stronger deblocking filter for luminance. Deblocking can implement a stronger deblocking filter for chrominance. VVC deblocking can implement a deblocking filter for sub-block boundaries. Deblocking can adapt deblocking decisions to small motion differences.
[0194] For example, deblocking filter parameters can be transmitted in the header based on the following logic. For example, PPS can be transmitted in the header based on the following logic: .
[0195] The `pps_deblocking_filter_control_present_flag` (e.g., equal to 1) can specify that a deblocking filter control syntax element exists in PPS. The `pps_deblocking_filter_control_present_flag` (e.g., equal to 0) can specify that a deblocking filter control syntax element does not exist in PPS and / or uses a value of 0 for deblocking. and t C The offset applies the deblocking filter to slices that reference PPS (e.g., all).
[0196] The `pps_deblocking_filter_override_enabled_flag` (e.g., equal to 1) specifies that deblocking behavior for images referencing PPS can be overridden at the image or slice level. The `pps_deblocking_filter_override_enabled_flag` (e.g., equal to 0) specifies that deblocking behavior for images referencing PPS is not overridden at the image or slice level. The value of `pps_deblocking_filter_override_enabled_flag` can be inferred (e.g., if it doesn't exist, it should be equal to 0).
[0197] The `pps_deblocking_filter_disabled_flag` (e.g., equal to 1) can specify that the deblocking filter is disabled for images referencing PPS, unless overridden by rendering PH or SH information for the image or slice. The `pps_deblocking_filter_disabled_flag` (e.g., equal to 0) can specify that the deblocking filter is enabled for images referencing PPS, unless overridden by rendering PH or SH information for the image or slice. For example, when the flag is not present, the value of `pps_deblocking_filter_disabled_flag` can be inferred (e.g., it should be equal to 0).
[0198] `pps_dbf_info_in_ph_flag` (e.g., equal to 1) can specify that deblocking filter information exists in the PH syntax structure, but not in the slice header referencing a PPS that does not include the PH syntax structure. `pps_dbf_info_in_ph_flag` (e.g., equal to 0) can specify that deblocking filter information does not exist in the PH syntax structure, but can exist in the slice header referencing a PPS. For example, when the flag is not present, the value of `pps_dbf_info_in_ph_flag` can be inferred (e.g., it should be equal to 0).
[0199] For example, pps_luma_beta_offset_div2 and pps_luma_tc_offset_div2 can specify the luminance component to be applied to slices referencing PPS. The default deblocking parameter offset and tC (e.g., divided by 2) are used, unless the default deblocking parameter offset is overridden by a deblocking parameter offset present in the image header and / or slice header of the PPS slice. The values of pps_luma_beta_offset_div2 and pps_luma_tc_offset_div2 can be within a range (e.g., from -12 to 12, inclusive). For example, their values can be inferred (e.g., to be equal to 0) when pps_luma_beta_offset_div2 and / or pps_luma_tc_offset_div2 do not exist.
[0200] For example, pps_cb_beta_offset_div2 and pps_cb_tc_offset_div2 can specify the Cb component to be applied to slices referencing PPS. The default deblocking parameter offset and tC (e.g., divided by 2), unless the default deblocking parameter offset is overridden by a deblocking parameter offset present in the image header and / or slice header of the PPS slice. The values of pps_cb_beta_offset_div2 and pps_cb_tc_offset_div2 can be within a range (e.g., in the range of -12 to 12, inclusive). For example, their values can be inferred (e.g., to be equal to pps_luma_beta_offset_div2 and pps_luma_tc_offset_div2 respectively) when pps_cb_beta_offset_div2 and / or pps_cb_tc_offset_div2 do not exist.
[0201] For example, pps_cr_beta_offset_div2 and pps_cr_tc_offset_div2 can specify the Cr component to be applied to slices referencing PPS. The default deblocking parameter offset and tC (e.g., divided by 2) are used, unless the default deblocking parameter offset is overridden by a deblocking parameter offset present in the image header and / or slice header of the PPS slice. The values of pps_cr_beta_offset_div2 and pps_cr_tc_offset_div2 can be within a range (e.g., from -12 to 12, inclusive). When the values of pps_cr_beta_offset_div2 and / or pps_cr_tc_offset_div2 do not exist, they can be inferred (e.g., to be equal to pps_luma_beta_offset_div2 and pps_luma_tc_offset_div2 respectively).
[0202] For example, the image header can be determined and / or transmitted based on the following logic: .
[0203] `ph_deblocking_params_present_flag` (e.g., equal to 1) specifies that deblocking parameters can exist in the PH syntax construct. `ph_deblocking_params_present_flag` (e.g., equal to 0) specifies that deblocking parameters do not exist in the PH syntax construct. For example, when `ph_deblocking_params_present_flag` does not exist, its value can be inferred (e.g., it should be equal to 0).
[0204] The `ph_deblocking_filter_disabled_flag` (e.g., equal to 1) can specify that the deblocking filter is disabled for the current image. The `ph_deblocking_filter_disabled_flag` (e.g., equal to 0) can specify that the deblocking filter is enabled for the current image.
[0205] ph_luma_beta_offset_div2 and ph_luma_tc_offset_div2 can specify the luminance component to be applied to the slice in the current image. The deblocking parameter offset of tC (e.g., divided by 2). The values of ph_luma_beta_offset_div2 and / or ph_luma_tc_offset_div2 can be within a range (e.g., in the range of -12 to 12, inclusive). For example, their values can be inferred when ph_luma_beta_offset_div2 and / or ph_luma_tc_offset_div2 do not exist (e.g., to be equal to pps_luma_beta_offset_div2 and pps_luma_tc_offset_div2 respectively).
[0206] ph_cb_beta_offset_div2 and ph_cb_tc_offset_div2 can specify the Cb component to be applied to the slice in the current image. And the deblocking parameter offset of tC (e.g., divided by 2). The values of ph_cb_beta_offset_div2 and / or ph_cb_tc_offset_div2 can be in a range (e.g., in the range of -12 to 12, inclusive).
[0207] ph_cr_beta_offset_div2 and ph_cr_tc_offset_div2 can specify the Cr component to be applied to the slice in the current image. And the deblocking parameter offset of tC (e.g., divided by 2). The values of ph_cr_beta_offset_div2 and / or ph_cr_tc_offset_div2 can be in a range (e.g., in the range of -12 to 12, inclusive).
[0208] For example, the slice header can be determined and / or transmitted based on the following logic: .
[0209] The `sh_deblocking_params_present_flag` (e.g., equal to 1) specifies that deblocking parameters can exist in the slice header. The `sh_deblocking_params_present_flag` (e.g., equal to 0) specifies that deblocking parameters do not exist in the slice header. For example, when `sh_deblocking_params_present_flag` does not exist, its value can be inferred (e.g., it should be equal to 0).
[0210] The `sh_deblocking_filter_disabled_flag` (e.g., equal to 1) can specify that the deblocking filter is disabled for the current slice. The `sh_deblocking_filter_disabled_flag` (e.g., equal to 0) can specify that the deblocking filter is enabled for the current slice.
[0211] sh_luma_beta_offset_div2 and sh_luma_tc_offset_div2 can specify the luma component to be applied to the current slice. The deblocking parameter offset of tC (e.g., divided by 2). The values of sh_luma_beta_offset_div2 and / or sh_luma_tc_offset_div2 can be within a range (e.g., in the range of -12 to 12, inclusive). For example, their values can be inferred when sh_luma_beta_offset_div2 and / or sh_luma_tc_offset_div2 do not exist (e.g., to be equal to ph_luma_beta_offset_div2 and ph_luma_tc_offset_div2 respectively).
[0212] sh_cb_beta_offset_div2 and sh_cb_tc_offset_div2 can specify the Cb component to be applied to the current slice. And the deblocking parameter offset of tC (e.g., divided by 2).
[0213] sh_cr_beta_offset_div2 and sh_cr_tc_offset_div2 can specify the Cr component to be applied to the current slice. And the deblocking parameter offset of tC (e.g., divided by 2). The values of sh_cr_beta_offset_div2 and / or sh_cr_tc_offset_div2 can be in a range (e.g., in the range of -12 to 12, inclusive).
[0214] The post-filtering order can be generated and signaled. Signaling the post-filtering order specifies the order (e.g., the correct order) used at the encoder and decoder for post-processing. The post-processing order can be stored in the header (e.g., in a special code). In some examples, the post-processing order can be stored in the slice header. In some examples, the post-processing order can be stored in the image header.
[0215] For example, exponential Golomb codes can be used to encode the post-processing order. Examples of exponential Golomb codes for post-filter order could be: .
[0216] In one example, the post-processing order could be as follows: FIG. 14 The realization described. For example... FIG. 14 As illustrated in the diagram, the post-processing sequence may include the computation of SAO, bilateral filter (BIF), and / or CCSAO. SAO can be applied (e.g., SAO can be applied before CCSAO). Each of the SAO offset and CCSAO offset can be applied, as described herein. ALF can be applied after CCSAO, as described herein.
[0217] In one example, the post-processing order could be as follows: FIG. 15 The realization described. For example... FIG. 15 As illustrated in the diagram, the SAO, bilateral filter (BIF), and / or CCSAO can be calculated. The CCSAO can then be applied (e.g., applied before the SAO). Each of the CCSAO offset and the SAO offset can be applied, as described herein. Finally, the ALF can be applied, as described herein.
[0218] For example, the post-processing order can be stored in a region header (e.g., a slice header) or an image header. On the encoder side, each post-processing order may generate encoding cost and distortion. A better cost / distortion tradeoff can provide a better (e.g., the best) order. A rate-distortion optimizer can perform optimizations.
[0219] This article provides examples of storing the processing order in the image header and / or slice header.
[0220] The processing order can be stored in the image header. In some instances, the luma post-filtering order can be similarly applied to chroma. In some examples, different orders may exist for luma and chroma. For example: .
[0221] For example, based on the following logic and / or syntax, the image header can indicate the processing order: .
[0222] The `rdo_postprocessing_order_picture()` function can return the order of brightness (e.g., the optimal order). The order of brightness (e.g., the optimal order) can be implemented as follows: .
[0223] The rate-distortion (RD) cost of a candidate block can be determined or calculated by the encoder. For example, the RD cost can be calculated according to equation (26): , Referring to equation (26), D can be the distortion of the current block and its coding parameters (e.g., L2, such as mean square error, peak signal-to-noise ratio (PSNR)). R can be the associated rate or cost (e.g., in terms of the number of coded bits). Lambda (QP) can be a Lagrangian parameter derived from the quantization parameter (QP). CostRD can be calculated for the (e.g., all) CTUs of the image.
[0224] The processing order can be stored in the slice header. Luminosity order can be applied to chroma. The processing order for luminosity can differ from the processing order for chroma. For example: .
[0225] For example, the slice header can indicate the processing order based on the following logic and / or syntax: .
[0226] The `rdo_postprocessing_order_slice()` function can be used to calculate or obtain the order of brightness (e.g., the optimal order). An example algorithm can be implemented as follows: .
[0227] The rate-distortion (RD) cost of a candidate block is a function that can be calculated by the encoder. The RD cost of a candidate block can be calculated, for example, according to equation (27): , Referring to equation (27), D can be the distortion of the current block with the current block coding parameters (e.g., L2, such as mean squared error, PSNR). R can be the associated rate or cost (e.g., in terms of the number of coded bits). Lambda(QP) can be the Lagrangian parameter derived from the quantization parameter QP. CostRD can be calculated for the slice's (e.g., all) CTUs.
[0228] In some examples, signaling at other levels can be performed. For instance, the post-processing order (e.g., postproc_order) can be signaled at the Picture Parameter Set (PPS) level. Encoder-side RDO can also be performed, for example, as in the case of signaling at the picture header level.
[0229] In some examples, the `postproc_order` syntax element can be inferred (e.g., as a default value) instead of being signaled. For instance, if (e.g., only) one postfilter is active, the `postproc_order` syntax element can be inferred as the default value instead of signaling it. Similarly, if the DBF is active along with another (e.g., only one other) postfilter, the `postproc_order` syntax element can be inferred as the default value instead of signaling it, where the DBF can be processed first.
[0230] In some examples, DBFs (e.g., if activated) can (e.g., always) be processed first. When DBFs are processed first, the number of combinations can be reduced, which saves signaling and encoder processing power.
[0231] In some examples, the syntax element `postproc_order_change` can be signaled before `postproc_order`. A `postproc_order_change` set to a first value (e.g., 0) can indicate the use of the default order (e.g., DBF, SAO, CCSAO, and ALF). Otherwise, `postproc_order` can be signaled to change the processing order. For example, if one or zero loop filters are activated (e.g., one or zero loop filters plus DBF), the value / indication of `postproc_order_change` can be inferred (e.g., 0).
[0232] Although the features and elements have been described above in specific combinations, those skilled in the art will understand that each feature or element can be used alone or in any combination with other features and elements. Furthermore, the methods described herein can be incorporated into a computer-readable medium to be implemented in a computer program, software, or firmware executed by a computer or processor. Examples of computer-readable media include electronic signals (transmitted via a wired or wireless connection) and computer-readable storage media. Examples of computer-readable storage media include, but are not limited to, read-only memory (ROM), random access memory (RAM), registers, cache memory, semiconductor memory devices, magnetic media (such as internal hard disks and removable disks), magneto-optical media, and optical media (such as CD-ROM discs and digital multifunction discs (DVDs)). The processor associated with the software can be used to implement a radio frequency transceiver used in a WTRU, UE, terminal, base station, RNC, or any host computer.
Claims
1. A video encoding device comprising a processor, the processor being configured to: Determine the order of post-filters associated with sub-regions of a region or image, where, The post-filter order is associated with at least two filters, and wherein the post-filter order indicates the processing order of the post-filters; and The order of the post-filters is sent in the region header or image header.
2. The video encoding device according to claim 1, wherein, The processor is further configured to: Generate encoding cost and distortion, wherein the encoding cost and distortion are related to the processing order of the post-filter.
3. The video encoding device according to any one of claims 1-2, wherein, The processor is further configured to determine the order of the luminance post-filter or the chrominance post-filter.
4. The video encoding device according to any one of claims 1-2, wherein, The processor is further configured to: Determine the order of the luminance post-filters; and The luminance post-filter order is applied to the chrominance post-filter order.
5. The video encoding device according to any one of claims 3-4, wherein, The order of the luminance post-filters is different from the order of the chrominance post-filters.
6. The video encoding device according to any one of claims 1-5, wherein, The processor is further configured to: When the post-filter sequence is sent to the region header, the rate distortion cost of the sub-regions of the region is determined.
7. The video encoding device according to any one of claims 1-5, wherein, The processor is further configured to: When the post-filter sequence is sent to the image header, the rate distortion cost of the sub-region of the image is determined.
8. The video encoding device according to any one of claims 6-7, wherein, The processor is further configured to: The rate distortion cost is determined based on the distortion of the current block with the encoding parameters of the current block, the associated rate or cost, and parameters derived from the quantization parameters.
9. The video encoding device according to any one of claims 1-6 or 8, wherein, The subregion is a coding tree unit (CTU), the region is a slice, and the region header is a slice header.
10. The video encoding device according to any one of claims 1-9, wherein, The at least two filters include Sample Adaptive Offset (SAO), Cross Component SAO (CCSAO), Adaptive Loop Filter, or Deblocking Filter.
11. The video encoding device according to any one of claims 1-10, wherein, The post-filter order indicates that cross-component sample adaptive offset (CCSAO) should be applied before sample adaptive offset (SAO).
12. A video decoding device comprising a processor, the processor being configured to: Determine the order of post-filters associated with sub-regions of a region or image, where, The post-filter order is associated with at least two filters, and wherein the post-filter order indicates the processing order of the post-filters; and The order of the post-filters is sent in the region header or image header.
13. The video decoding device according to claim 12, wherein, The processor is further configured to: Generate encoding cost and distortion, wherein the encoding cost and distortion are related to the processing order of the post-filter.
14. The video decoding device according to any one of claims 12-13, wherein, The processor is further configured to determine the order of the luminance post-filter or the chrominance post-filter.
15. The video decoding device according to any one of claims 12-13, wherein, The processor is further configured to: Determine the order of the luminance post-filters; and The luminance post-filter order is applied to the chrominance post-filter order.
16. The video decoding device according to any one of claims 14-15, wherein, The order of the luminance post-filters is different from the order of the chrominance post-filters.
17. The video decoding device according to any one of claims 12-16, wherein, The processor is further configured to: When the post-filter sequence is sent to the region header, the rate distortion cost of the sub-regions of the region is determined.
18. The video decoding device according to any one of claims 12-16, wherein, The processor is further configured to: When the post-filter sequence is sent to the image header, the rate distortion cost of the sub-region of the image is determined.
19. The video decoding device according to any one of claims 17-18, wherein, The processor is further configured to: The rate distortion cost is determined based on the distortion of the current block with the encoding parameters of the current block, the associated rate or cost, and parameters derived from the quantization parameters.
20. The video decoding device according to any one of claims 12-17 or 19, wherein, The subregion is a coding tree unit (CTU), the region is a slice, and the region header is a slice header.
21. The video decoding device according to any one of claims 12-20, wherein, The at least two filters include Sample Adaptive Offset (SAO), Cross Component SAO (CCSAO), Adaptive Loop Filter, or Deblocking Filter.
22. The video decoding device according to any one of claims 12-21, wherein, The post-filter order indicates that cross-component sample adaptive offset (CCSAO) should be applied before sample adaptive offset (SAO).
23. A method for video encoding, comprising: Determine the order of post-filters associated with sub-regions of a region or image, where, The post-filter order is associated with at least two filters, and wherein the post-filter order indicates the post-processing order; and The order of the post-filters is sent in the region header or image header.
24. The video encoding method according to claim 23, further comprising: Generate encoding cost and distortion, among which, The encoding cost and the distortion are related to the processing order of the post-filter.
25. The video coding method according to any one of claims 23-24, further comprising determining the order of the luminance post-filter or the chrominance post-filter.
26. The video coding method according to any one of claims 23-24, further comprising: Determine the order of the luminance post-filters; and The luminance post-filter order is applied to the chrominance post-filter order.
27. The video coding method according to any one of claims 25-26, wherein, The order of the luminance post-filters is different from the order of the chrominance post-filters.
28. The video coding method according to any one of claims 23-27, further comprising: When the post-filter sequence is sent to the region header, the rate distortion cost of the sub-regions of the region is determined.
29. The video coding method according to any one of claims 23-27, further comprising: When the post-filter sequence is sent to the image header, the rate distortion cost of the sub-region of the image is determined.
30. The video coding method according to any one of claims 28-29, further comprising: The rate distortion cost is determined based on the distortion of the current block with the encoding parameters of the current block, the associated rate or cost, and parameters derived from the quantization parameters.
31. The video coding method according to any one of claims 23-28 or 30, wherein, The subregion is a coding tree unit (CTU), the region is a slice, and the region header is a slice header.
32. The video coding method according to any one of claims 23-31, wherein, The at least two filters include Sample Adaptive Offset (SAO), Cross Component SAO (CCSAO), Adaptive Loop Filter, or Deblocking Filter.
33. The video coding method according to any one of claims 23-32, wherein, The post-filter order indicates that cross-component sample adaptive offset (CCSAO) should be applied before sample adaptive offset (SAO).
34. A method for video decoding, comprising: Determine the order of post-filters associated with sub-regions of a region or image, where, The post-filter order is associated with at least two filters, and wherein the post-filter order indicates the processing order of the post-filters; and The order of the post-filters is sent in the region header or image header.
35. The video decoding method according to claim 34, further comprising: Generate encoding cost and distortion, among which, The encoding cost and the distortion are related to the processing order of the post-filter.
36. The video decoding method according to any one of claims 34-35, further comprising determining the order of the luminance post-filter or the chrominance post-filter.
37. The video decoding method according to any one of claims 34-35, further comprising: Determine the order of the luminance post-filters; and The luminance post-filter order is applied to the chrominance post-filter order.
38. The video decoding method according to any one of claims 36-37, wherein, The order of the luminance post-filters is different from the order of the chrominance post-filters.
39. The video decoding method according to any one of claims 34-38, further comprising: When the post-filter sequence is sent to the region header, the rate distortion cost of the sub-regions of the region is determined.
40. The video decoding method according to any one of claims 34-38, further comprising: When the post-filter sequence is sent to the image header, the rate distortion cost of the sub-region of the image is determined.
41. The video decoding method according to any one of claims 39-40, further comprising: The rate distortion cost is determined based on the distortion of the current block with the encoding parameters of the current block, the associated rate or cost, and parameters derived from the quantization parameters.
42. The video decoding method according to any one of claims 34-39 or 41, wherein, The subregion is a coding tree unit (CTU), the region is a slice, and the region header is a slice header.
43. The video decoding device according to any one of claims 34-42, wherein, The at least two filters include Sample Adaptive Offset (SAO), Cross Component SAO (CCSAO), Adaptive Loop Filter, or Deblocking Filter.
44. The video decoding device according to any one of claims 34-43, wherein, The post-filter order indicates that cross-component sample adaptive offset (CCSAO) should be applied before sample adaptive offset (SAO).
45. A video encoding device comprising a processor, wherein, The processor is configured to implement the steps of the method according to any one of claims 23-33.
46. A video decoding device comprising a processor, wherein, The processor is configured to implement the steps of the method according to any one of claims 34-44.
47. A computer program product stored on a non-transitory computer-readable medium and comprising program code instructions for implementing the steps of the method according to any one of claims 23-33 or 34-44.
48. Video data, comprising information representing encoded blocks generated by the method according to any one of claims 23-33.