Quantization center shift due to rate gradient
By shifting the quantization center based on the rate gradient, the system optimizes video coding efficiency by adjusting quantization exponents, improving bitrate management and quality adaptation in video encoding and decoding.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- INTERDIGITALCE PATENT HLDG SAS
- Filing Date
- 2024-04-05
- Publication Date
- 2026-05-19
AI Technical Summary
Existing video coding systems face challenges in optimizing quantization processes to efficiently manage bitrate and quality in video encoding and decoding, particularly in adapting to varying entropy rates.
The system shifts the quantization center by a gradient of the rate, using a reciprocal of the quantization exponent to adjust conversion coefficients, which can be variable, fixed, or dynamically generated, to improve quantized transformation coefficients.
This approach enhances the efficiency of video coding by optimizing bitrate allocation and quality, addressing the limitations of traditional quantization methods in adapting to entropy rate variations.
Smart Images

Figure 2026515612000001_ABST
Abstract
Description
[Technical Field]
[0001] Cross-reference of related applications
[0001] This application claims the benefits of European Patent Application No. 23305513.6, filed on 7 April 2023, and European Patent Application No. 23305916.1, filed on 8 June 2023, the contents of which are incorporated herein by reference. [Background technology]
[0002]
[0002] A video coding system may be used to compress a digital video signal, for example, to reduce the storage and / or transmission bandwidth required for such a signal. Video coding systems may include, for example, block-based, wavelet-based, and / or object-based systems. [Overview of the project]
[0003]
[0003] The system, method, and means are configured to shift the quantization center by a gradient of the rate. A video coding device (e.g., a video encoding and / or video decoding device) may obtain a quantization exponent. The device may shift the quantization exponent based on a quantity that is the reciprocal of the quantization exponent. The device may apply the shifted quantization exponent to a conversion coefficient to generate an adjusted conversion coefficient.
[0004]
[0004] The device may determine that the quantity that is the reciprocal of the quantization exponent includes the gradient of the entropy rate. The device may determine that the quantity that is the reciprocal of the quantization exponent is based on a lookup table. The quantity may be a variable quantity, a fixed quantity, or a dynamically generated true gradient. The transformation coefficient may be a quantized transformation coefficient. The transformation coefficient may be an inversely quantized transformation coefficient. The quantization exponent may include an inversely quantized exponent. The adjusted coefficient may include an inversely quantized transformation coefficient. [Brief explanation of the drawing]
[0005] Brief Description of the Drawings [Figure 1A]
[0005] FIG. 1 is a system diagram illustrating an exemplary communication system in which one or more of the disclosed embodiments may be implemented. [Figure 1B]
[0006] FIG. 2 is a system diagram illustrating an exemplary wireless transmit / receive unit (WTRU) that may be used within the communication system illustrated in FIG. 1A, according to one embodiment. [Figure 1C]
[0007] FIG. 3 is a system diagram illustrating an exemplary radio access network (RAN) and an exemplary core network (CN) that may be used within the communication system illustrated in FIG. 1A, according to one embodiment. [Figure 1D]
[0008] FIG. 4 is a system diagram illustrating a further exemplary RAN and a further exemplary CN that may be used within the communication system illustrated in FIG. 1A, according to one embodiment. [Figure 2]
[0009] FIG. 5 illustrates an exemplary video encoder. [Figure 3]
[0010] FIG. 6 illustrates an exemplary video decoder. [Figure 4]
[0011] FIG. 7 illustrates an example of a system in which various aspects and examples may be implemented. [Figure 5]
[0012] FIG. 8 illustrates an exemplary video encoder. [Figure 6]
[0013] FIG. 9 illustrates an exemplary video decoder.
Best Mode for Carrying Out the Invention
[0006] Detailed Description
[0014] A more detailed understanding may be obtained from the following description given by way of example in conjunction with the accompanying drawings.
[0007] <00Figure 1A illustrates an exemplary communication system 100 in which one or more disclosed embodiments may be implemented. The communication system 100 may be a multiple access system that provides content such as voice, data, video, messaging, and broadcast to multiple wireless users. The communication system 100 may enable multiple wireless users to access such content through the sharing of system resources, including wireless bandwidth. For example, the communication system 100 may employ one or more channel access methods, such as code division multiple access (CDMA), time division multiple access (TDMA), frequency division multiple access (FDMA), orthogonal FDMA (OFDMA), single-carrier FDMA (SC-FDMA), zero-tail unique-word DFT-Spread OFDM (ZT UW DTS-s OFDM), unique-word OFDM (UW-OFDM), resource block filter OFDM, filter bank multicarrier (FBMC), and similar methods.
[0008]
[0016] As shown in Figure 1A, the communication system 100 may include wireless transmit / receive units (WTRUs) 102a, 102b, 102c, 102d, RAN 104 / 113, CN 106 / 115, public switched telephone network (PSTN) 108, the Internet 110, and other networks 112, but it should be understood that the disclosed embodiments intend any number of WTRUs, base stations, networks, and / or network elements. Each of the WTRUs 102a, 102b, 102c, and 102d may be any type of device configured to operate and / or communicate in a wireless environment. For example, WTRU102a, 102b, 102c, and 102d (all of which may be called “stations” and / or “STAs”) may be configured to transmit and / or receive radio signals and may include user equipment (UEs), mobile stations, fixed or mobile subscriber units, subscriber base units, pagers, mobile phones, personal digital assistants (PDAs), smartphones, laptops, netbooks, personal computers, wireless sensors, hotspots or Mi-Fi devices, Internet of Things (IoT) devices, watches or other wearables, head-mounted displays (HMDs), vehicles, drones, medical devices and applications (e.g., remote surgery), industrial devices and applications (e.g., robots and / or other wireless devices operating in industrial and / or automated processing chain settings), consumer electronics devices, and devices operating on commercial and / or industrial wireless networks. WTRU102a, 102b, 102c, and 102d may all be interchangeably referred to as UEs.
[0009]
[0017] The communication system 100 may also include base stations 114a and / or base stations 114b. Each of the base stations 114a and 114b may be any type of device configured to wirelessly interface with at least one of the WTRUs 102a, 102b, 102c, and 102d to facilitate access to one or more communication networks such as CN 106 / 115, the Internet 110, and / or other networks 112. For example, base stations 114a and 114b may be a base transceiver station (BTS), node B, eNode B, home node B, home eNode B, gNB, NR node B, site controller, access point (AP), wireless router, etc. Although base stations 114a and 114b are depicted as single elements, it should be understood that base stations 114a and 114b may include any number of interconnected base stations and / or network elements.
[0010]
[0018] Base station 114a may be part of RAN 104 / 113, which may also include other base stations and / or network elements (not shown) (such as a base station controller (BSC), a radio network controller (RNC), and relay nodes). Base stations 114a and / or base stations 114b may be configured to transmit and / or receive radio signals on one or more carrier frequencies which may be called cells (not shown). These frequencies may exist in the licensed spectrum, the unlicensed spectrum, or a combination of the licensed and unlicensed spectrum. A cell may provide coverage of a radio service to a particular geographic area which may be relatively fixed or may change over time. A cell may be further divided into cell sectors. For example, a cell associated with base station 114a may be divided into three sectors. Thus, in one embodiment, base station 114a may include three transceivers, i.e., one transceiver per sector of the cell. In one embodiment, the base station 114a may employ multiple-input multiple output (MIMO) technology and utilize multiple transceivers for each sector of the cell. For example, beamforming may be used to transmit and / or receive signals in a desired spatial direction.
[0011]
[0019] Base stations 114a and 114b may communicate with one or more WTRUs 102a, 102b, 102c, and 102d via a radio interface 116, the radio interface 116 may be any suitable radio communication link (e.g., radio frequency (RF), microwave, centimeter wave, micrometer wave, infrared (IR), ultraviolet (UV), visible light, etc.). The radio interface 116 may be established using any suitable radio access technology (RAT).
[0012]
[0020] More specifically, as described above, the communication system 100 may be a multiple access system and may employ one or more channel access schemes such as CDMA, TDMA, FDMA, OFDMA, SC-FDMA, and the like. For example, base stations 114a and WTRU 102a, 102b, 102c within RAN 104 / 113 may implement radio technologies such as Universal Mobile Telecommunications System (UMTS) Terrestrial Radio Access (UTRA) that can establish radio interfaces 115 / 116 / 117 using wideband CDMA (WCDMA). WCDMA may include communication protocols such as High-Speed Packet Access (HSPA) and / or Advanced HSPA (HSPA+). HSPA may include High-Speed Downlink (DL) Packet Access (HSDPA) and / or High-Speed UL Packet Access (HSUPA).
[0013]
[0021] In one embodiment, base stations 114a and WTRUs 102a, 102b, and 102c may implement radio technologies such as Evolved UMTS Terrestrial Radio Access (E-UTRA), which can establish a radio interface 116 using Long Term Evolution (LTE) and / or LTE-Advanced (LTE-A) and / or LTE-Advanced Pro (LTE-A Pro).
[0014]
[0022] In one embodiment, base stations 114a and WTRUs 102a, 102b, and 102c may implement radio technologies such as NR radio access, which can establish a radio interface 116 using New Radio (NR).
[0015]
[0023] In one embodiment, base station 114a and WTRU 102a, 102b, 102c may implement multiple radio access technologies. For example, base station 114a and WTRU 102a, 102b, 102c may implement LTE radio access and NR radio access together, for example, using the dual connectivity (DC) principle. Thus, the radio interface used by WTRU 102a, 102b, 102c may be characterized by transmissions between multiple types of radio access technologies and / or multiple types of base stations (e.g., eNB and gNB).
[0016]
[0024] In other embodiments, base stations 114a and WTRUs 102a, 102b, and 102c may implement wireless technologies such as IEEE 802.11 (i.e., Wireless Fidelity (WiFi)), IEEE 802.16 (i.e., Worldwide Interoperability for Microwave Access (WiMAX)), CDMA2000, CDMA2000 1X, CDMA2000 EV-DO, Provisional Standard 2000 (IS-2000), Provisional Standard 95 (IS-95), Provisional Standard 856 (IS-856), Global System for Mobile Communications (GSM), Enhanced Data rates for GSM Evolution (EDGE), and GSM EDGE (GERAN).
[0017]
[0025] The base station 114b in Figure 1A may be, for example, a wireless router, home node B, home eNode B, or access point, and may utilize any suitable RAT to facilitate wireless connectivity in local areas such as offices, homes, vehicles, campuses, industrial facilities, aerial walkways (e.g., for use by drones), roads, and the like. In one embodiment, the base station 114b and WTRU 102c, 102d may implement wireless technologies such as IEEE 802.11 to establish a wireless local area network (WLAN). In one embodiment, the base station 114b and WTRU 102c, 102d may implement wireless technologies such as IEEE 802.15 to establish a wireless personal area network (WPAN). In yet another embodiment, base stations 114b and WTRUs 102c, 102d may establish picocells or femtocells using cellular-based RATs (e.g., WCDMA, CDMA2000, GSM, LTE, LTE-A, LTE-A Pro, NR, etc.). As shown in Figure 1A, base station 114b may have a direct connection to the internet 110. Therefore, base station 114b may not need to access the internet 110 via CN 106 / 115.
[0018]
[0026] RAN104 / 113 may communicate with CN106 / 115, which may be any type of network configured to provide voice, data, applications, and / or Voice over Internet Protocol (VoIP) services to one or more of WTRU102a, 102b, 102c, and 102d. The data may have various quality of service (QoS) requirements, such as different throughput requirements, latency requirements, error tolerance requirements, reliability requirements, data throughput requirements, mobility requirements, and so on. CN106 / 115 may provide call control, billing services, mobile location-based services, prepaid calls, internet connectivity, video distribution, and / or perform high-level security functions such as user authentication. Although not shown in Figure 1A, it should be understood that RAN104 / 113 and / or CN106 / 115 may communicate directly or indirectly with other RANs employing the same or different RATs as RAN104 / 113. For example, in addition to connecting to RAN104 / 113 which may be using NR radio technology, CN106 / 115 may also communicate with another RAN (not shown) employing GSM, UMTS, CDMA2000, WiMAX, E-UTRA, or WiFi radio technology.
[0019]
[0027] CN106 / 115 may also function as a gateway for WTRU102a, 102b, 102c, and 102d to access PSTN108, the Internet 110, and / or other networks 112. PSTN108 may include a circuit-switched telephone network providing plain old telephone service (POTS). The Internet 110 may include a global system of interconnected computer networks and devices using common communication protocols such as the transmission control protocol (TCP), user datagram protocol (UDP), and / or Internet protocol (IP) within the TCP / IP Internet Protocol suite. Network 112 may include wired and / or wireless networks owned and / or operated by other service providers. For example, network 112 may include another CN connected to one or more RANs that may employ the same RAT as RAN104 / 113 or a different RAT.
[0020]
[0028] Some or all of the WTRUs 102a, 102b, 102c, and 102d in the communication system 100 may include multimode functionality (for example, WTRUs 102a, 102b, 102c, and 102d may include multiple transceivers for communicating with different radio networks via different radio links). For example, WTRU 102c shown in Figure 1A may be configured to communicate with base station 114a, which may employ cellular-based radio technology, and base station 114b, which may employ IEEE 802 radio technology.
[0021]
[0029] Figure 1B is a system diagram showing an exemplary WTRU 102. As shown in Figure 1B, the WTRU 102 may comprise, among other things, a processor 118, a transceiver 120, a transmit / receive element 122, a speaker / microphone 124, a keypad 126, a display / touchpad 128, non-removable memory 130, removable memory 132, a power supply 134, a global positioning system (GPS) chipset 136, and / or other peripherals 138. It should be understood that the WTRU 102 may include any partial combination of the above elements while maintaining consistency with one embodiment.
[0022]
[0030] The processor 118 may be a general-purpose processor, a dedicated processor, a conventional processor, a digital signal processor (DSP), multiple microprocessors, one or more microprocessors associated with a DSP core, a controller, a microcontroller, an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) circuit, any other type of integrated circuit (IC), a state machine, etc. The processor 118 may perform signal coding, data processing, power control, input / output processing, and / or any other functions that enable the WTRU 102 to operate in a wireless environment. The processor 118 may be coupled to a transceiver 120 which may be coupled to a transmit / receive element 122. Although Figure 1B depicts the processor 118 and the transceiver 120 as separate components, it should be understood that the processor 118 and the transceiver 120 may be integrated together in an electronic package or chip.
[0023]
[0031] The transmit / receive element 122 may be configured to transmit signals to or receive signals from a base station (e.g., base station 114a) via the radio interface 116. For example, in one embodiment, the transmit / receive element 122 may be an antenna configured to transmit and / or receive RF signals. In one embodiment, the transmit / receive element 122 may be an emitter / detector configured to transmit and / or receive, for example, IR, UV, or visible light signals. In yet another embodiment, the transmit / receive element 122 may be configured to transmit and / or receive both RF signals and optical signals. It should be understood that the transmit / receive element 122 may be configured to transmit and / or receive any combination of radio signals.
[0024]
[0032] In Figure 1B, the transmit / receive element 122 is shown as a single element, but the WTRU 102 may have any number of transmit / receive elements 122. More specifically, the WTRU 102 may employ MIMO technology. Thus, in one embodiment, the WTRU 102 may include two or more transmit / receive elements 122 (e.g., multiple antennas) for transmitting and receiving radio signals via the radio interface 116.
[0025]
[0033] The transceiver 120 may be configured to modulate the signal to be transmitted by the transmit / receive element 122 and to demodulate the signal received by the transmit / receive element 122. As described above, the WTRU 102 may have multimode capabilities. Therefore, the transceiver 120 may include multiple transceivers to enable the WTRU 102 to communicate via multiple RATs, such as NR and IEEE 802.11.
[0026]
[0034] The processor 118 of the WTRU102 may be coupled to a speaker / microphone 124, a keypad 126, and / or a display / touchpad 128 (e.g., a liquid crystal display (LCD) display unit or an organic light-emitting diode (OLED) display unit) and may receive user input data from them. The processor 118 may also output user data to the speaker / microphone 124, the keypad 126, and / or the display / touchpad 128. In addition, the processor 118 may access information from any type of suitable memory, such as non-removable memory 130 and / or removable memory 132, and store data in such memory. The non-removable memory 130 may include random-access memory (RAM), read-only memory (ROM), a hard disk, or any other type of memory storage device. The removable memory 132 may include a subscriber identity module (SIM) card, a memory stick, a secure digital (SD) memory card, and the like. In other embodiments, the processor 118 may access information from memory not physically located on the WTRU 102, such as on a server or home computer (not shown), and store data in such memory.
[0027]
[0035] The processor 118 may receive power from the power supply 134 and may be configured to distribute and / or control power to other components within the WTRU 102. The power supply 134 may be any suitable device for supplying power to the WTRU 102. For example, the power supply 134 may be one or more dry cell batteries (e.g., nickel-cadmium (NiCd), nickel-zinc (NiZn), nickel metal hydride (NiMH), lithium-ion (Li-ion), etc.), a solar cell, a fuel cell, etc.
[0028]
[0036] The processor 118 may also be coupled to a GPS chipset 136 which can be configured to provide location information (e.g., longitude and latitude) regarding the current location of the WTRU 102. In addition to or instead of information from the GPS chipset 136, the WTRU 102 may receive location information from base stations (e.g., base stations 114a, 114b) via the radio interface 116 and / or determine its location based on the timing of signals received from two or more nearby base stations. It should be understood that the WTRU 102 may acquire location information by any suitable location determination method while maintaining consistency with one embodiment.
[0029]
[0037] The processor 118 may be further coupled to other peripherals 138, which may include one or more software modules and / or hardware modules that provide additional features, functions and / or wired or wireless connectivity. Examples of peripherals 138 include accelerometers, e-compasses, satellite transceivers, digital cameras (for photography and / or video), universal serial bus (USB) ports, vibration devices, television transceivers, hands-free headsets, Bluetooth® modules, frequency modulated (FM) wireless units, digital music players, media players, video game player modules, internet browsers, virtual reality and / or augmented reality (VR / AR) devices, activity trackers, and the like. The peripheral device 138 may include one or more sensors, which may be one or more of the following: gyroscope, accelerometer, Hall effect sensor, magnetometer, compass sensor, proximity sensor, temperature sensor, time sensor, geolocation sensor, altimeter, light sensor, touch sensor, magnetometer, barometer, gesture sensor, biometric sensor, and / or humidity sensor.
[0030]
[0038] WTRU102 may include a full-duplex radio (for example, one in which the transmission and reception of some or all of the signals associated with a particular subframe for both UL (e.g., for transmission) and downlink (e.g., for reception) may be in parallel and / or simultaneous. The full-duplex radio may include an interference management unit for reducing and / or substantially eliminating self-interference either through hardware (e.g., chokes) or signal processing via a processor (e.g., a separate processor (not shown) or processor 118). In one embodiment, WTRU102 may include a half-duplex radio that transmits and receives some or all of the signals (for example, one in which the signals associated with a particular subframe for either UL (e.g., for transmission) or downlink (e.g., for reception)).
[0031]
[0039] Figure 1C is a system diagram illustrating RAN104 and CN106 according to one embodiment. As described above, RAN104 employs E-UTRA wireless technology and can communicate with WTRU102a, 102b, and 102c via wireless interface 116. RAN104 can also communicate with CN106.
[0032]
[0040] RAN104 may include eNode-B160a, 160b, and 160c, but it should be understood that RAN104 may include any number of eNode-B while maintaining consistency with one embodiment. Each eNode-B160a, 160b, and 160c may include one or more transceivers for communicating with WTRU102a, 102b, and 102c via the radio interface 116. In one embodiment, eNode-B160a, 160b, and 160c may implement MIMO technology. For example, eNode-B160a may use multiple antennas to transmit radio signals to and / or receive radio signals from WTRU102a.
[0033]
[0041] Each of the eNode-B160a, 160b, and 160c may be associated with a specific cell (not shown) and may be configured to handle wireless resource management decisions, handover decisions, user scheduling in UL and / or DL, etc. As shown in Figure 1C, the eNode-B160a, 160b, and 160c may communicate with each other via the X2 interface.
[0034]
[0042] The CN106 shown in Figure 1C may include a mobility management entity (MME) 162, a serving gateway (SGW) 164, and a packet data network (PDN) gateway (or PGW) 166. Although each of the above elements is depicted as part of CN106, it should be understood that any of these elements may be owned and / or operated by an entity other than the CN operator.
[0035]
[0043] The MME162 can be connected to each of the eNode-B162a, 162b, and 162c within RAN 104 via the S1 interface and can function as a control node. For example, the MME162 may be responsible for user authentication of WTRU102a, 102b, and 102c, bearer activation / deactivation, and selecting a specific serving gateway during the initial attachment of WTRU102a, 102b, and 102c. The MME162 may provide control plane functionality for switching between RAN 104 and other RANs (not shown) employing other radio technologies such as GSM and / or WCDMA.
[0036]
[0044] The SGW164 can be connected to each of the eNode B160a, 160b, and 160c within RAN104 via the S1 interface. The SGW164 can generally route and forward user data packets to and from WTRU102a, 102b, and 102c. The SGW164 can perform other functions such as anchoring the user plane during eNode B handovers, triggering paging when DL data is available for WTRU102a, 102b, and 102c, managing and remembering the context of WTRU102a, 102b, and 102c, and similar functions.
[0037]
[0045] SGW164 can be connected to PGW166, thereby providing WTRU102a, 102b, and 102c with access to packet-switched networks such as the Internet 110, facilitating communication between WTRU102a, 102b, and 102c and IP-enabled devices.
[0038]
[0046] CN106 can facilitate communication with other networks. For example, CN106 can provide WTRU102a, 102b, and 102c with access to a circuit-switched network such as PSTN108 to facilitate communication between WTRU102a, 102b, and 102c and conventional land communication devices. For example, CN106 may include or communicate with an IP gateway (e.g., an IP multimedia subsystem (IMS (IP multimedia subsystem)) server) that acts as an interface between CN106 and PSTN108. In addition, CN106 may provide WTRU102a, 102b, and 102c with access to other networks 112, which may include other wired and / or wireless networks owned and / or operated by other service providers.
[0039]
[0047] In Figures 1A to 1D, the WTRU is described as a wireless terminal, but in certain representative embodiments, it is intended that such a terminal may have access to a wired communication interface with a communication network (e.g., temporarily or permanently).
[0040]
[0048] In a typical embodiment, the other network 112 may be a WLAN.
[0041]
[0049] A WLAN in Infrastructure Basic Service Set (BSS) mode may have access points (APs) for the BSS and one or more stations (STAs) associated with those APs. APs may have access to or interfaces with a Distribution System (DS) or another type of wired / wireless network that carries traffic to and from the BSS. Traffic originating outside the BSS and destined for an STA may arrive via an AP and be delivered to the STA. Traffic originating from an STA to a destination outside the BSS may be sent to an AP and delivered to its respective destination. Traffic between STAs within the BSS may be transmitted via an AP; for example, a source STA may send traffic to an AP, which then delivers the traffic to a destination STA. Traffic between STAs within the BSS may be considered and / or referred to as peer-to-peer traffic. Peer-to-peer traffic may be transmitted (e.g., directly) between a source STA and a destination STA using a direct link setup (DLS). In certain representative embodiments, the DLS may use 802.11e DLS or 802.11z tunneled DLS (TDLS (tunneled DLS)). A WLAN using Independent BSS (IBSS (Independent BSS)) mode may not have APs, and STAs within or using IBSS (e.g., all STAs) may communicate directly with each other. The IBSS communication mode may sometimes be referred to as “ad hoc” communication mode in this specification.
[0042]
[0050] When using the 802.11ac infrastructure operating mode or a similar operating mode, an AP may transmit beacons on a fixed channel, such as a primary channel. The primary channel may be of a fixed width (e.g., a wideband of 20 MHz) or a width dynamically set via signal transmission. The primary channel may be the operating channel of the BSS and may be used by an STA to establish a connection with the AP. In certain typical embodiments, for example in an 802.11 system, Carrier Sense Multiple Access with Collision Avoidance (CSMA / CA) may be implemented. In the case of CSMA / CA, an STA, including the AP (e.g., any STA), may sense the primary channel. If the primary channel is sensed / detected by a particular STA and / or determined to be busy, that particular STA may backoff. In a given BSS, one STA (e.g., only one station) may transmit at any given time.
[0043]
[0051] A high-throughput (HT) STA can use a 40MHz wide channel for communication, for example, by combining a primary 20MHz channel with adjacent or non-adjacent 20MHz channels to form a 40MHz wide channel.
[0044]
[0052] Very High Throughput (VHT) STAs can support 20MHz, 40MHz, 80MHz, and / or 160MHz wide channels. 40MHz and / or 80MHz channels can be formed by combining consecutive 20MHz channels. 160MHz channels can be formed by combining eight consecutive 20MHz channels, or by combining two non-consecutive 80MHz channels, which may be called an 80+80 configuration. In the 80+80 configuration, the channel-coded data can pass through a segment parser that can split the data into two streams. Inverse Fast Fourier Transform (IFFT) and time-domain processing can be performed separately for each stream. The streams can be mapped to two 80MHz channels, and the data can be transmitted by a transmitting STA. The receiver of the receiving STA can reverse the operation described above for the 80+80 configuration and transmit the combined data to Medium Access Control (MAC).
[0045]
[0053] Operating modes below 1 GHz are supported in 802.11af and 802.11ah. Channel operating bandwidth and carrier are reduced in 802.11af and 802.11ah compared to those used in 802.11n and 802.11ac. 802.11af supports bandwidths of 5 MHz, 10 MHz, and 20 MHz in the TV White Space (TVWS) spectrum, while 802.11ah supports bandwidths of 1 MHz, 2 MHz, 4 MHz, 8 MHz, and 16 MHz using the non-TVWS spectrum. According to a typical embodiment, 802.11ah may support meter-type control / machine-type communications, such as MTC devices in a macro coverage area. MTC devices may have limited functionality, including support for specific and / or limited bandwidths (e.g., support only). MTC devices may include batteries with battery life exceeding a threshold (e.g., to maintain very long battery life).
[0046]
[0054] WLAN systems that can support multiple channels and channel bandwidths, such as 802.11n, 802.11ac, 802.11af, and 802.11ah, include a channel that can be designated as the primary channel. The primary channel may have a bandwidth equal to the maximum common operating bandwidth supported by all STAs in the BSS. The bandwidth of the primary channel may be set and / or limited by the STA that supports the minimum bandwidth operating mode from among all STAs operating in the BSS. In the 802.11ah example, even if the AP and other STAs in the BSS support operating modes of 2MHz, 4MHz, 8MHz, 16MHz and / or other channel bandwidths, the primary channel may be 1MHz wide for an STA (e.g., an MTC type device) that supports 1MHz mode (e.g., only 1MHz mode). Carrier sensing and / or network allocation vector (NAV) settings may depend on the status of the primary channel. For example, if an STA (which only supports 1MHz operating mode) has its primary channel busy transmitting to an AP, the entire available frequency band may be considered busy, even though a large portion of the frequency band remains idle and could be available.
[0047]
[0055] In the United States, the available frequency band for 802.11ah is 902MHz to 928MHz. In South Korea, the available frequency band is 917.5MHz to 923.5MHz. In Japan, the available frequency band is 916.5MHz to 927.5MHz. The total available bandwidth for 802.11ah is 6MHz to 26MHz, depending on the country code.
[0048]
[0056] Figure 1D is a system diagram illustrating RAN113 and CN115 according to one embodiment. As described above, RAN113 employs NR radio technology and can communicate with WTRU102a, 102b, and 102c via the radio interface 116. RAN113 can also communicate with CN115.
[0049]
[0057] Device RAN113 may include gNB180a, 180b, and 180c, but it will be understood that RAN113 may include any number of gNBs while maintaining consistency with the embodiment. gNB180a, 180b, and 180c may each include one or more transceivers for communicating with WTRU102a, 102b, and 102c via the radio interface 116. In one embodiment, gNB180a, 180b, and 180c may implement MIMO technology. For example, gNB180a and 180b may use beamforming to transmit signals to and / or receive signals from gNB180a, 180b, and 180c. Thus, for example, gNB180a may use multiple antennas to transmit radio signals to and / or receive radio signals from WTRU102a. In one embodiment, gNB180a, 180b, and 180c may implement carrier aggregation technology. For example, gNB180a may transmit multiple component carriers to WTRU102a (not shown). A subset of these component carriers may be on the unauthorized spectrum, and the remaining component carriers may be on the authorized spectrum. In one embodiment, gNB180a, 180b, and 180c may implement coordinated multi-point (CoMP) technology. For example, WTRU102a may receive coordinated transmissions from gNB180a and gNB180b (and / or gNB180c).
[0050]
[0058] WTRU102a, 102b, and 102c may communicate with gNB180a, 180b, and 180c using transmissions associated with scalable neurology. For example, OFDM symbol intervals and / or OFDM subcarrier intervals may vary for different transmissions, different cells, and / or different parts of the radio transmission spectrum. WTRU102a, 102b, and 102c may communicate with gNB180a, 180b, and 180c using subframes or transmission time intervals (TTIs) of varying or scalable lengths (e.g., containing varying numbers of OFDM symbols and / or lasting for varying absolute times).
[0051]
[0059] gNB180a, 180b, and 180c can be configured to communicate with WTRU102a, 102b, and 102c in standalone and / or non-standalone configurations. In a standalone configuration, WTRU102a, 102b, and 102c can communicate with gNB180a, 180b, and 180c without accessing other RANs (e.g., eNode-B160a, 160b, and 160c). In a standalone configuration, WTRU102a, 102b, and 102c can use one or more gNB180a, 180b, and 180c as mobility anchor points. In a standalone configuration, WTRU102a, 102b, and 102c can communicate with gNB180a, 180b, and 180c using signals in an unauthorized band. In a non-standalone configuration, WTRU102a, 102b, and 102c can communicate / connect with gNB180a, 180b, and 180c, while also communicating / connecting with other RANs such as eNode-B160a, 160b, and 160c. For example, WTRU102a, 102b, and 102c can implement DC principles for substantially simultaneous communication with one or more gNB180a, 180b, and 180c and one or more eNode-B160a, 160b, and 160c. In a non-standalone configuration, eNode-B160a, 160b, and 160c can function as mobility anchors for WTRU102a, 102b, and 102c, and gNB180a, 180b, and 180c can provide additional coverage and / or throughput to service WTRU102a, 102b, and 102c.
[0052]
[0060] Each of the gNB180a, 180b, and 180c may be associated with a specific cell (not shown) and may be configured to address wireless resource management decisions, handover decisions, user scheduling in UL and / or DL, support for network slicing, dual connectivity, interaction between NR and E-UTRA, routing of user plane data to User Plane Functions (UPFs) 184a and 184b, and routing of control plane information to Access and Mobility Management Functions (AMFs) 182a and 182b. As shown in Figure 1D, the gNB180a, 180b, and 180c may communicate with each other via the Xn interface.
[0053]
[0061] The CN115 shown in Figure 1D may include at least one AMF182a, 182b, at least one UPF184a, 184b, at least one Session Management Function (SMF)183a, 183b, and optionally a Data Network (DN)185a, 185b. Although each of the above elements is depicted as part of the CN115, it should be understood that any of these elements may be owned and / or operated by an entity other than the CN operator.
[0054]
[0062] AMF182a and 182b can be connected to one or more gNB180a, 180b, and 180c within RAN113 via the N2 interface and can function as control nodes. For example, AMF182a and 182b may be responsible for user authentication of WTRU102a, 102b, and 102c, support for network slicing (e.g., handling different PDU sessions with different requirements), selection of specific SMF183a and 183b, management of registration areas, termination of NAS signaling, and mobility management. Network slicing may be used by AMF182a and 182b to customize CN support for WTRU102a, 102b, and 102c based on the type of service being utilized by WTRU102a, 102b, and 102c. For example, different network slices may be established for different use cases, such as services that rely on ultra-reliable low latency (URLLC) access, services that rely on enhanced massive mobile broadband (eMBB) access, services for machine type communication (MTC) access, and / or similar. The AMF162 may provide control plane functionality for switching between RAN113 and other RANs (not shown) employing other radio technologies such as LTE, LTE-A, LTE-A Pro, and / or non-3GPP access technologies such as WiFi.
[0055]
[0063] SMF183a and 183b can be connected to AMF182a and 182b in CN115 via the N11 interface. SMF183a and 183b can also be connected to UPF184a and 184b in CN115 via the N4 interface. SMF183a and 183b can select and control UPF184a and 184b and configure the routing of traffic passing through UPF184a and 184b. SMF183a and 183b can perform other functions such as managing and assigning UE IP addresses, managing PDU sessions, controlling policy enforcement and QoS, and providing downlink data notifications. PDU session types can be IP-based, non-IP-based, Ethernet-based, etc.
[0056]
[0064] UPF184a and 184b can be connected to one or more gNB180a, 180b, and 180c in RAN113 via the N3 interface, thereby providing WTRU102a, 102b, and 102c with access to packet-switched networks such as the Internet 110, facilitating communication between WTRU102a, 102b, and 102c and IP-enabled devices. UPF184 and 184b can perform other functions such as packet routing and forwarding, enforcement of user plane policies, support for multi-homed PDU sessions, handling user plane QoS, buffering downlink packets, providing mobility anchoring, and similar functions.
[0057]
[0065] CN115 can facilitate communication with other networks. For example, CN115 may include, or communicate with, an IP gateway (e.g., an IP multimedia subsystem (IMS) server) that functions as an interface between CN115 and PSTN108. In addition, CN115 may provide WTRU102a, 102b, 102c with access to other networks 112, which may include other wired and / or wireless networks owned and / or operated by other service providers. In one embodiment, WTRU102a, 102b, 102c may be connected to local data networks (DN) 185a, 185b via UPF184a, 184b through an N3 interface to UPF184a, 184b and an N6 interface between UPF184a, 184b and DN185a, 185b.
[0058]
[0066] Considering Figures 1A to 1D and their corresponding descriptions, one or more of the functions described herein with respect to one or more of the WTRU102a to d, base stations 114a to b, eNode-B160a to c, MME162, SGW164, PGW166, gNB180a to c, AMF182a to b, UPF184a to b, SMF183a to b, DN185a to b, and / or any other devices described herein, one or more of the functions described herein may be performed by one or more emulation devices (not shown). An emulation device may be one or more devices configured to emulate one or more of the functions described herein. For example, an emulation device may be used to test other devices and / or simulate network and / or WTRU functions.
[0059]
[0067] Emulation devices may be designed to perform one or more tests on other devices in a laboratory and / or operator network environment. For example, one or more emulation devices may perform one or more or all functions when fully or partially implemented and / or deployed as part of a wired and / or wireless network to test other devices in a communications network. One or more emulation devices may perform one or more or all functions when temporarily implemented / deployed as part of a wired and / or wireless network. Emulation devices may be directly coupled to another device for testing purposes and / or perform tests using wireless communications.
[0060]
[0068] One or more emulation devices may perform one or more functions, including all functions, without being implemented / deployed as part of a wired and / or wireless communication network. For example, an emulation device may be used in a test scenario in a test laboratory and / or an undeployed (e.g., test) wired and / or wireless communication network to perform testing of one or more components. One or more emulation devices may be test equipment. Wireless communication via direct RF coupling and / or RF circuitry (e.g., which may include one or more antennas) may be used by an emulation device to transmit and / or receive data.
[0061]
[0069] This application describes various embodiments, including tools, features, examples, models, and methods. Many of these embodiments are described in detail, often in a manner that may seem restrictive, at least in order to illustrate their individual characteristics. However, this is intended to clarify the description and not to limit the application or scope of those embodiments. In fact, all of the different embodiments can be combined and interchangeable to provide further embodiments. Moreover, these embodiments can also be combined and interchangeable with embodiments described in prior applications.
[0062]
[0070] The embodiments described and contemplated herein can be implemented in many different forms. Figures 5 and 6 described herein may provide examples, but other examples are also contemplated. The considerations in Figures 5 and 6 do not limit the scope of implementation forms. At least one of the embodiments generally relates to video encoding and decoding, and at least one other embodiment generally relates to transmitting a generated or encoded bitstream. These embodiments and other embodiments can be implemented as a computer-readable storage medium storing instructions for encoding or decoding video data according to any of the described methods, apparatus, and / or a computer-readable storage medium storing a bitstream generated according to any of the described methods.
[0063]
[0071] In this application, the terms “reconstructed” and “decoded” may be used interchangeably, the terms “pixel” and “sample” may be used interchangeably, and the terms “image,” “picture,” and “frame” may be used interchangeably.
[0064]
[0072] Various methods are described herein, each of which includes one or more steps or actions to achieve the described method. Unless a particular order of steps or actions is required for the proper operation of the method, the order and / or use of any particular steps and / or actions may be modified or combined. In addition, terms such as “first,” “second,” etc., may be used in various examples to modify elements, components, steps, actions, etc. (e.g., “first decryption” and “second decryption”). The use of such terms does not implicitly imply an ordering of the modified actions unless specifically required. Thus, in this example, the first decryption does not need to be performed before the second decryption, but may be performed, for example, before the second decryption, during the second decryption, or during a period overlapping with the second decryption.
[0065]
[0073] As shown in Figures 2 and 3, various methods and other embodiments described herein may be used to modify modules of the video encoder 200 and decoder 300, for example, the decoding module. Furthermore, the subject matter disclosed herein may be applied to any type, format, or version of video encoding, whether existing or future, whether described in standards or recommendations, and whether they are described in standards or recommendations. Unless otherwise indicated or technically excluded, the embodiments described herein may be used individually or in combination.
[0066]
[0074] Various numerical values are used in the examples described in this application. These and other specific values are for illustrative purposes only, and the embodiments described are not limited to these specific values.
[0067]
[0075] Figure 2 shows an exemplary video encoder. While variations of the exemplary encoder 200 are intended, encoder 200 is described below for clarity and not all expected variations are described.
[0068]
[0076] Before encoding, the video sequence may undergo pre-encoding processing (201), such as applying a color conversion to the input color picture (e.g., from RGB4:4:4 to YCbCr4:2:0) or performing a remapping of the input picture components to obtain a signal distribution more resistant to compression (e.g., using histogram equalization of one of the color components). Metadata may be associated with the pre-processing and appended to the bitstream.
[0069]
[0077] In encoder 200, the picture is encoded by encoder elements as described below. The picture to be encoded is divided into units of coding units (CUs) (202) and processed. Each unit is encoded using either intra-mode or inter-mode, for example. When a unit is encoded in intra-mode, it performs intra-prediction (260). In inter-mode, motion estimation (275) and compensation (270) are performed. The encoder decides whether to use intra-mode or inter-mode to encode a unit (205), and indicates the intra / inter decision by, for example, a prediction mode flag. For example, the prediction residual is calculated by subtracting the prediction block from the original image block (210).
[0070]
[0078] The predicted residual is then transformed (225) and quantized (230). The quantized transformation coefficients, as well as other syntactic elements such as motion vectors and picture segmentation information, are entropy coded to output a bitstream (245). The encoder can skip the transformation and apply quantization directly to the untransformed residual signal. The encoder can bypass both the transformation and quantization, i.e., the residual is coded directly without applying either the transformation or quantization process.
[0071]
[0079] The encoder decodes the encoded blocks to provide a reference for further prediction. The quantized transformation coefficients are inversely quantized (240) and inversely transformed (250) to decode the prediction residuals. The decoded prediction residuals and prediction blocks are combined (255) to reconstruct the image blocks. For example, an in-loop filter (265) is applied to the reconstructed picture to perform deblocking / SAO (Sample Adaptive Offset) / ALF (Adaptive Loop Filter) filtering to reduce encoding artifacts. The filtered image is stored in a reference picture buffer (280).
[0072]
[0080] Figure 3 shows an example of a video decoder. In the exemplary decoder 300, the bitstream is decoded by the decoder elements as described below. The video decoder 300 generally performs a decoding path that is the reverse of the encoding path described in Figure 2. The encoder 200 also generally performs video decoding as part of the encoding of video data.
[0073]
[0081] In particular, the decoder input includes a video bitstream that may be generated by the video encoder 200. The bitstream is first entropy-decoded to obtain transformation coefficients, prediction modes, motion vectors, and other coded information (330). Picture segmentation information indicates how the picture is segmented. Thus, the decoder may segment the picture according to the decoded picture segmentation information (335). The transformation coefficients are inversely quantized (340) and inversely transformed (350) to decode the prediction residuals. The decoded prediction residuals and prediction blocks are combined (355) to reconstruct an image block. The prediction block may be obtained from intra-prediction (360) or motion-compensated prediction (i.e., inter-prediction) (375) (370). An in-loop filter (365) is applied to the reconstructed image. The filtered image is stored in a reference picture buffer (380). In some cases, for a given picture, the contents of the reference picture buffer 380 on the decoder 300 side may be identical to the contents of the reference picture buffer 280 on the encoder 200 side (for the same picture).
[0074]
[0082] The decoded picture may undergo further post-decoded processing (385), such as reverse color transformation (e.g., conversion from YCbCr 4:2:0 to RGB 4:4:4), or reverse remapping, which performs the reverse of the remapping process performed in pre-encoded processing (201). Post-decoded processing may use metadata derived in pre-encoded processing and signaled in a bitstream. In some examples, the decoded image (e.g., after applying the in-loop filter (365) and / or after post-decoded processing (385) if post-decoded processing is used) may be sent to a display device for rendering to the user.
[0075]
[0083] Figure 4 shows an example of a system in which the various embodiments and examples described herein may be implemented. System 400 may be embodied as a device including the various components described below and configured to perform one or more of the embodiments described herein. Examples of such devices include, but are not limited to, various electronic devices such as personal computers, laptop computers, smartphones, tablet computers, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected consumer electronics, and servers. The elements of System 400 may be embodied individually or in combination as a single integrated circuit (IC), multiple ICs, and / or separate components. For example, in at least one example, the processing elements and encoder / decoder elements of System 400 are distributed across multiple ICs and / or separate components. In various examples, System 400 is communicably coupled to one or more other systems or other electronic devices, for example, via a communication bus or via dedicated input and / or output ports. In various examples, System 400 is configured to implement one or more of the embodiments described herein.
[0076]
[0084] System 400 includes, for example, at least one processor 410 configured to execute loaded instructions in order to implement various embodiments described herein. The processor 410 may include embedded memory, input / output interfaces, and various other circuits well known in the art. System 400 includes at least one memory 420 (e.g., a volatile memory device and / or a non-volatile memory device). System 400 includes a storage device 440, which may include non-volatile memory and / or volatile memory, including, but not limited to, electrically erasable programmable read-only memory (EEPROM), read-only memory (ROM), programmable read-only memory (PROM), random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), flash memory, magnetic disk drives, and / or optical disk drives. The storage device 440 may, in non-limiting examples, include internal storage devices, mounted storage devices (including removable and non-removable storage devices), and / or network-accessible storage devices.
[0077]
[0085] System 400 includes, for example, an encoder / decoder module 430 configured to process data and provide encoded or decoded video, the encoder / decoder module 430 of which may include its own processor and memory. The encoder / decoder module 430 represents a module that may be included in the device to perform encoding and / or decoding functions. As is well known, the device may include one or both of the encoding and decoding modules. In addition, the encoder / decoder module 430 may be implemented as a separate element of System 400, or it may be incorporated into the processor 410 as a combination of hardware and software, as is known to those skilled in the art.
[0078]
[0086] Program code loaded onto the processor 410 or encoder / decoder 430 to perform the various embodiments described herein may be stored in the storage device 440 and subsequently loaded into memory 420 for execution by the processor 410. According to various examples, one or more of the processor 410, memory 420, storage device 440, and encoder / decoder module 430 may store one or more of various items during the performance of the processes described herein. Such stored items include, but are not limited to, input video, decoded video or a portion of decoded video, bitstreams, matrices, variables, and intermediate or final results from the processing of equations, formulas, operations, and operational logic.
[0079]
[0087] In some examples, the internal memory of the processor 410 and / or the encoder / decoder module 430 is used to store instructions and provide working memory for the processing required during encoding or decoding. However, in other examples, memory outside the processing device (for example, the processing device may be either the processor 410 or the encoder / decoder module 430) is used for one or more of these functions. The external memory may be memory 420 and / or storage device 440, such as dynamic volatile memory and / or non-volatile flash memory. In some examples, external non-volatile flash memory is used to store, for example, the television's operating system. In at least one example, high-speed external dynamic volatile memory, such as RAM, is used as working memory for video encoding and decoding operations.
[0080]
[0088] Inputs to the elements of system 400 may be provided via various input devices as shown in block 445. Such input devices include, but are not limited to, (i) a radio frequency (RF) portion for receiving RF signals transmitted by broadcast by a broadcasting station, (ii) a component (COMP) input terminal (or a set of COMP input terminals), (iii) a universal serial bus (USB) input terminal, and / or (iv) a high-definition multimedia interface (HDMI) input terminal. Another example, although not shown in Figure 4, is composite video.
[0081]
[0089] In various examples, the input device of block 445 has associated input processing elements, as is well known in the art. For example, the RF portion may be associated with elements suitable for (i) selecting a desired frequency (also called selecting a signal, or band-limiting a signal to a frequency in a certain band), (ii) down-converting the selected signal, (iii) again band-limiting to a narrower frequency band in order to select a signal frequency band that may be called a channel in a particular example, (iv) demodulating the down-converted and band-limited signal, (v) performing error correction, and / or (vi) demultiplexing to select a desired data packet stream. The RF portion of various examples includes one or more elements for performing these functions, e.g., frequency selectors, signal selectors, band limiters, channel selectors, filters, downconverters, demodulators, error correctors, and demultiplexers. The RF portion may include tuners that perform various of these functions, including, for example, down-converting a received signal to a lower frequency (e.g., an intermediate frequency or a frequency near the baseband) or to the baseband. In one example set-top box, the RF section and its associated input processing elements receive an RF signal transmitted via a wired (e.g., cable) medium and perform frequency selection by filtering, down-converting, and re-filtering to a desired frequency band. Various examples involve rearranging the order of the elements described above (and others), removing some of these elements, and / or adding other elements that perform similar or different functions. Adding elements can include inserting elements between existing elements, such as inserting an amplifier and an analog-to-digital converter. In various examples, the RF section includes an antenna.
[0082]
[0090] The USB and / or HDMI terminals may include their respective interface processors for connecting the system 400 to other electronic devices via USB and / or HDMI connections. It should be understood that various aspects of input processing, such as Reed-Solomon error correction, may be implemented, for example, in a separate input processing IC or within the processor 410 as needed. Similarly, aspects of USB or HDMI interface processing may be implemented, for example, in a separate interface IC or within the processor 410 as needed. Demodulated, error-corrected, and demultiplexed streams are provided to various processing elements, including, for example, the processor 410 and an encoder / decoder 430, which work in conjunction with memory and storage elements to process the data stream as needed for presentation on an output device.
[0083]
[0091] Various elements of system 400 may be housed within an integrated housing. Within the integrated housing, the various elements can be interconnected and data can be transmitted between them using an appropriate connection configuration 425, such as an internal bus known in the art, including an Inter-IC (I2C) bus, wiring, and a printed circuit board.
[0084]
[0092] System 400 includes a communication interface 450 that enables communication with other devices via a communication channel 460. The communication interface 450 may include, but is not limited to, a transceiver configured to transmit and receive data via the communication channel 460. The communication interface 450 may also include, but is not limited to, a modem or a network card, and the communication channel 460 may be implemented, for example, within a wired and / or wireless medium.
[0085]
[0093] In various examples, data is streamed to system 400 or otherwise provided using a wireless network such as a Wi-Fi network, e.g., IEEE 802.11 (IEEE stands for Institute of Electrical and Electronics Engineers). In these examples, the Wi-Fi signal is received via a communication channel 460 and a communication interface 450 adapted for Wi-Fi communication. In these examples, communication channel 460 is typically connected to an access point or router that provides access to an external network, including the Internet, to enable streaming applications and other over-the-top communications. Other examples provide the streamed data to system 400 using a set-top box that distributes data via an HDMI connection on input block 445. Yet another example provides the streamed data to system 400 using an RF connection on input block 445. As shown above, various examples provide data in non-streaming manners. In addition, various examples use wireless networks other than Wi-Fi, e.g., cellular networks or Bluetooth® networks.
[0086]
[0094] System 400 can provide output signals to various output devices, including a display 475, a speaker 485, and other peripheral devices 495. Various examples of the display 475 include, for example, one or more of a touchscreen display, an organic light-emitting diode (OLED) display, a curved display, and / or a foldable display. The display 475 may be for a television, tablet, laptop, mobile phone, or other device. The display 475 may also be integrated with other components (for example, in a smartphone) or separate (for example, an external monitor for a laptop). Other peripheral devices 495, in various examples, include one or more of a standalone digital video disc (or digital multipurpose disc) (both terms refer to DVD), a disc player, a stereo system, and / or a lighting system. Various examples use one or more peripheral devices 495 that provide functions based on the output of System 400. For example, a disc player performs the function of playing back the output of System 400.
[0087]
[0095] In various examples, control signals are transmitted between the system 400 and the display 475, speaker 485, or other peripheral devices 495 using signal transmission such as AV.Link, Consumer Electronics Control (CEC), or other communication protocols that enable control between devices with or without user intervention. Output devices may be communicably coupled to the system 400 via dedicated connectors through their respective interfaces 470, 480, and 490. Alternatively, output devices may be connected to the system 400 via the communication interface 450 using the communication channel 460. The display 475 and speaker 485 may be integrated into a single unit with other components of the system 400 in an electronic device such as a television. In various examples, the display interface 470 includes a display driver, such as a timing controller (TCon) chip.
[0088]
[0096] For example, if the RF portion of input 445 is part of a separate set-top box, the display 475 and speaker 485 can, alternatively, be separated from one or more of the other components. In various examples where the display 475 and speaker 485 are external components, the output signal may be provided via a dedicated output connection, such as an HDMI port, a USB port, or a COMP output. These examples may be implemented by computer software or hardware implemented by the processor 410, or by a combination of hardware and software. In a non-limiting example, these examples may be implemented by one or more integrated circuits. The memory 420 may be of any type appropriate for the technical environment and, in a non-limiting example, may be implemented using any suitable data storage technology such as optical memory devices, magnetic memory devices, semiconductor-based memory devices, fixed memory, and removable memory. The processor 410 may be of any type appropriate for the technical environment and, in a non-limiting example, may include one or more of a microprocessor, a general-purpose computer, a dedicated computer, and a processor based on a multi-core architecture.
[0089]
[0097] Various implementations include decoding. As used in this application, “decoding” may encompass all or part of the process performed on a received encoded sequence to produce, for example, a final output suitable for display. In various examples, such a process includes one or more processes typically performed by a decoder, e.g., entropy decoding, inverse quantization, inverse transformation, and differential decoding. In various examples, such a process also, or alternatively, includes processes performed by a decoder in the various implementations described in this application, e.g., obtaining a quantization exponent, shifting the quantization exponent based on a quantity that is the reciprocal of the quantization exponent, and applying the shifted quantization exponent to the transformation coefficient to produce a tuned transformation coefficient.
[0090]
[0098] As further examples, in one example, “decoding” refers only to entropy decoding; in another example, “decoding” refers only to differential decoding; and in yet another example, “decoding” refers to a combination of entropy decoding and differential decoding. Whether the expression “decoding process” is intended to refer specifically to a subset of operations or to the broader decoding process as a whole will become clear from the context of the specific description and should be well understood by those skilled in the art.
[0091]
[0099] Various implementations include encoding. Similar to what is discussed above with respect to "decoding," as used in this application, "encoding" may encompass all or part of the processes performed on an input video sequence to generate an encoded bitstream. In various examples, such processes typically include one or more processes performed by an encoder, e.g., splitting, differential coding, transformation, quantization, and entropy coding. In various examples, such processes also, or alternatively, include processes performed by an encoder in the various implementations described in this application, e.g., obtaining a quantization exponent, shifting the quantization exponent based on a quantity that is the reciprocal of the quantization exponent, and applying the shifted quantization exponent to the transformation coefficient to generate a tuned transformation coefficient.
[0092]
[0100] For further examples, in one instance, “encoding” refers only to entropy coding; in another instance, “encoding” refers only to differential coding; and in yet another instance, “encoding” refers to a combination of differential and entropy coding. Whether the expression “encoding process” is intended to refer specifically to a subset of operations or to the broader encoding process as a whole will become clear from the context of the specific description and should be well understood by those skilled in the art.
[0093]
[0101] Please note that the syntactic elements used in this specification are descriptive terms. Therefore, they do not preclude the use of other syntactic element names or functions.
[0094]
[0102] When a diagram is presented as a flow chart, it should be understood that the diagram also provides a block diagram of the corresponding device. Similarly, when a diagram is presented as a block diagram, it should be understood that the diagram also provides a flow chart of the corresponding method / process.
[0095]
[0103] The implementations and embodiments described herein may be implemented, for example, in the form of methods or processes, apparatus, software programs, data streams, or signals. Even if a single implementation is discussed only in the context of a single implementation (for example, only as a method), the implementation of the features discussed may also be implemented in other forms (for example, apparatus or programs). Apparatus may be implemented, for example, in the form of appropriate hardware, software, and firmware. A method may be implemented, for example, in a processor, where processor generally refers to a processing device including, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device. A processor also includes communication devices such as, for example, computers, mobile phones, portable / personal digital assistants ("PDAs"), and other devices that facilitate the communication of information between end users.
[0096]
[0104] References to “one example” or “one implementation” or “one implementation,” and other variations thereof, mean that the specific features, structures, characteristics, etc. described in relation to that example are included in at least one example. Therefore, the phrases “one example” or “one example” or “in one implementation” or “in one implementation,” and any other variations appearing in various places throughout this application, do not necessarily all refer to the same example.
[0097]
[0105] In addition, this application may refer to “determining” various types of information. Determining information may include, for example, one or more of the following: estimating information, calculating information, predicting information, or retrieving information from memory. Acquiring may include receiving, retrieving, constructing, generating, and / or determining.
[0098]
[0106] Furthermore, this application may refer to “accessing” various types of information. Accessing information may include, for example, receiving information, retrieving information (e.g., from memory), storing information, moving information, copying information, computing information, determining information, predicting information, or estimating information.
[0099]
[0107] In addition, this application may refer to "receiving" various types of information. Receiving is intended to be a broad term, similar to "accessing." Receiving information may include, for example, accessing information or retrieving information (for example, from memory) one or more of these. Furthermore, "receiving" is typically involved in some way during operations such as, for example, storing information, processing information, transmitting information, moving information, copying information, erasing information, calculating information, determining information, predicting information, or estimating information.
[0100]
[0108] Please understand that the use of any of the following " / ", "and / or", and "at least one of" is intended to encompass the selection of only the first enumerated option (A), or only the second enumerated option (B), or both options (A and B), for example, in the cases of "A / B", "A and / or B", and "at least one of A and B". As further examples, in the cases of "A, B, and / or C" and "at least one of A, B, and C", such phrases are intended to encompass the selection of only the first enumerated option (A), or only the second enumerated option (B), or only the third enumerated option (C), or only the first and second enumerated options (A and B), or only the first and third enumerated options (A and C), or only the second and third enumerated options (B and C), or all three options (A, B, and C). This can be extended to a number of items listed, as will be obvious to those skilled in the art in this and related fields.
[0101]
[0109] Furthermore, as used herein, the term “signal” refers, in particular, to indicating something to the corresponding decoder. An encoder signal may include, for example, a signal indicating whether the encoder has chosen an implicit decision, such as one or more adjustment values, or an explicit decision. In this way, in one example, the same parameter is used on both the encoder and decoder sides. Therefore, for example, the encoder can transmit a specific parameter to the decoder so that the decoder can use the same specific parameter (explicit signaling). Conversely, if the decoder already has a specific parameter and other parameters, signaling may be used without transmission, simply allowing the decoder to know and select the specific parameter (implicit signaling). Bit saving is achieved in various examples by avoiding the transmission of any actual function. It should be understood that signaling can be achieved in various ways. For example, in various examples, one or more syntax elements, flags, etc., are used to signal information to the corresponding decoder. The above concerns the verb form of the word "signal," but the word "signal" may also be used as a noun in this specification.
[0102]
[0110] As will be apparent to those skilled in the art, various implementations can generate a variety of signals formatted to carry information that can be stored or transmitted, for example. The information may include, for example, instructions for performing a method, or data generated by one of the implementations described. For example, a signal may be formatted to carry the bitstream of the example described. Such a signal may be formatted, for example, as an electromagnetic wave (e.g., using the radio frequency portion of the spectrum) or as a baseband signal. Formatting may include, for example, encoding a data stream and modulating a carrier with the encoded data stream. The information carried by the signal may be, for example, analog or digital information. The signal may be transmitted through a variety of different wired or wireless links, as is well known. The signal may be stored on a processor-readable medium, accessed from a processor-readable medium, or received.
[0103]
[0111] This specification provides numerous examples. The features of the examples may be provided individually or in any combination across various categories and types of claims. Furthermore, the examples may include, individually or in any combination across various categories and types of claims, one or more of the features, devices, or embodiments described herein. For example, the features described herein may be implemented in the form of a bitstream or signal containing information generated as described herein. This information may enable a decoder to decode the bitstream, and the encoder, bitstream, and / or decoder may follow any of the described embodiments. For example, the features described herein may be implemented by creating and / or transmitting and / or receiving and / or decoding a bitstream or signal. For example, the features described herein may be implemented as a method, process, apparatus, medium for storing instructions, medium for storing data, or signal. For example, the features described herein may be implemented by a television, set-top box, mobile phone, tablet, or other electronic device performing decoding. Televisions, set-top boxes, mobile phones, tablets, or other electronic devices may display the resulting image (e.g., an image from residual reconstruction of a video bitstream) (e.g., using a monitor, screen, or other type of display). Televisions, set-top boxes, mobile phones, tablets, or other electronic devices may receive and decode a signal containing an encoded image.
[0104]
[0112] Quantization can be achieved based on rate distortion optimization (RDO). Transformation coefficients can be quantized to encode the coefficients into a bitstream with a given rate-distortion tradeoff. The quantization process considers both quantization error (e.g., distortion) and coding cost (e.g., rate) objectives, and can find values that satisfy both objectives. In video compression, two techniques can be used to find values that satisfy both objectives. These may be trellis-coded quantization (TCQ) (also known as dependent quantization (DQ)) and rate distortion optimized quantization (RDOQ).
[0105]
[0113] TCQ uses two (e.g., different) uniform scalar quantizations, and can describe when the first or second quantizer is used by elegantly designed states and transitions between states. The selection of quantizers and their quantization exponents for input coefficients can be chosen according to rate and distortion objectives along with the designed states and transitions. The designed states and transitions can affect the search space (e.g., limiting it), and better possibilities can be explored (e.g., and ignored) to find a good trade-off between coding complexity and compression performance.
[0106]
[0114] RDOQ can use a single uniform quantization. RDOQ can find an assignment (e.g., another) that can increase quantization error and decrease coding cost (e.g., compared to a fixed quantization exponential assignment to coefficients (nearest grid center)), and obtain a rate-distortion tradeoff (e.g., an acceptable rate-distortion tradeoff).
[0107]
[0115] Quantization maps can be used in RDO-based quantizers. Uniform scalar quantization maps (e.g., their quantization boundaries and reconstruction centers) can be scalar quantizers for uniform sources at bitrates and for high bitrates for other sources when fixed-length coding (e.g., for distortion purposes) is considered. Uniform scalar quantizers can be used in lossy video compression. A series of studies may aim to design quantization maps for transform blocks in rate-distortion optimization (RDO) based quantization steps, which may signal the quantization map to be used by increasing the rate and introducing new coding complexity in order to find a transform block-specific quantization map.
[0108]
[0116] When a quantizer is operating under a distortion purpose, its optimality may be affected if a rate purpose is added. The quantizer used may be a uniform scalar quantizer. When a rate purpose is added, the optimality of the quantizer used may be affected. The quantization map may be adjusted to mitigate its impact on encoder / decoder complexity.
[0109]
[0117] To obtain a quantization map that satisfies rate and distortion objectives, the centers of the quantization grid (e.g., reconstruction points) can be adjusted. RDO-based quantization processes can be considered multi-objective, one example being that satisfy the Karush-Kuhn-Tucker (KKT) condition. Using the KKT condition, the gradient with respect to the quantization error (e.g., distortion) objective can be negative (in the opposite direction) of the gradient with respect to the coding cost (e.g., rate) objective. By changing the direction, one can be used in place of the other.
[0110]
[0118] According to first-order gradient-based optimization, for example, if one step is taken along the negative gradient with respect to the objective, the objective is improved. To reduce distortion, the inversely quantized coefficients can be shifted along the negative direction of the gradient with respect to the distortion. The gradient with respect to distortion does not necessarily have to exist on the decoder side. If the gradient with respect to distortion is in the negative direction of the gradient with respect to rate, then the gradient with respect to rate can be used on the decoder side.
[0111]
[0119] If an RDO-based quantizer does not have an explicit differentiable rate prediction (e.g., the rate can be measured using an arithmetic encoder), a gradient surrogate with respect to the rate may be used. The gradient step size can be adjusted universally or it may be specific to the image and / or transformation block.
[0112]
[0120] Image / video compression methods may use RDO-based quantization, which can be generalized using equation (1).
number
[0113]
[0121] In equation (1), x ∈ R n is the real coefficient of the length n to be quantized, where y∈Z n This can be a quantization exponent defined on a discrete set of reconstruction points. Quantizer Q(.) and inverse quantizer function Q -1 Using (.), the exponent and reconstruction are given by y=Q(x) and
number
[0114]
[0122] RDO-based quantization can be an unconstrained multi-objective optimization problem where the objectives can be the minimum rate and the minimum reconstruction error. The following can actually show the characteristics of an example of an unconstrained multi-objective optimization problem.
[0115]
[0123] An example of multi-objective optimization can exist when it satisfies the Karush-Kuhn-Tucker (KKT) conditions. More specifically, in the case of unconstrained multi-objective optimization, for the objective in Equation (2): [Number] (where α i ≥ 0 and Σ i α i = 1, and L i is the objective function), y* can be optimal when it satisfies the following, as can be seen in Equation (3): Σ I α i ∇ y L i (y*) = 0 (3)
[0116]
[0124] In the following example, the conditions that an RDO-based quantizer should have can be presented.
[0117]
[0125] RDO-based quantization adjusted using the trade-off of λ can satisfy a set of conditions for a given quantization and inverse quantization function when the following conditions are met, as can be seen in Equation (4). ∇y (R(y))=-λ∇ y (D(x,Q -1 (y))) (4) α1=1, α2=1, L1(y):=R(y), and L2(y):=λD(x,Q -1 When (y)), equation (1) can be rewritten in the form of equation (2). RDO-based quantization can be an unconstrained multi-objective optimization. RDO-based quantization can satisfy equation (3) using the previous notation, and equation (3) is ∇ y (R(y))+λ∇ y (D(x,Q -1 This can be rewritten by (y)))=0, which can be equal to equation (4).
[0118]
[0126] If one of the exemplary RDO-based quantizers having a given quantization and inverse quantization function can find a solution (y*), then the gradient of the quantization exponent with respect to distortion may be negative than the gradient of the quantization exponent with respect to rate, and the gradients with respect to distortion and rate may be used interchangeably.
[0119]
[0127] Some RDO-based quantization methods, such as TCQ and RDOQ, or methods in video codecs, can minimize rate distortion and find a solution (e.g., an optimal solution). This is not always the case for both TCQ and RDOQ. These methods may not thoroughly explore the (e.g., entire) space and may ignore candidate solutions.
[0120]
[0128] The features described herein can be associated with gradient transformations. In first-order gradient-based optimization, for example, if we take one step in the negative direction of its gradient with respect to the objective function, the objective function may become smaller (e.g., smaller). For RDO-based quantization, y←y-ρ∇ y (D(x,Q -1 (y)))(where ρ∈R + If the quantization exponent y is shifted by the gradient with respect to the strain (where is the step size), then the strain D(x,Q) -1(y) can be smaller. Distortion cannot be known unless the original input data is known (e.g., explicitly). By using the examples described herein, a negative gradient with respect to rate can be used so that it is the gradient with respect to distortion. The quantization exponent can be shifted within the loop of the dequantization block of the encoder and the dequantization block of the decoder by equation (5). y←y+ρ*∇ y (R(y)) (5)
[0121]
[0129] In equation (5), the quantization exponent is the (e.g., small) step size ρ*∇ y (R(y)) can be shifted in a direction that increases the rate. If a closed form of the rate function (entropy model) is known and differentiable, its gradient can be obtained (e.g., explicitly) and used in equation (5). The exponent can be shifted (e.g., explicitly) by the theoretical direction. The best step size ρ* can be adjusted through a set of validations and used as a fixed value that can be universally applied. In one example, the best step size may be variable. The best step size can be determined for a picture or transformation block and signaled to the decoder. If a closed form of the rate is not known, it can be calculated experimentally by an arithmetic encoder. The rate can be modeled using known basic, manually designed functions or neural networks to have differentiable rate predictions.
[0122]
[0130] To generalize the unavailability of differentiable entropy model (e.g., rate) cases, a differentiable entropy model case can be modeled by a parametric function f(.|θ) parameterized by θ, along with input of available information of the transformation block, thereby,
number
number
[0123]
[0131] In some cases, differentiable entropy models (e.g., rate) may not be available and can be modeled by linear models using quantized exponents, thereby R(y i )≒f(y i ,|a,b)=a|y i |+b. The rate of the quantization exponent can be independent of others and can be increased by the increment of the absolute value of the quantization exponent. R(y i )≒a|y i Since |+b,
number
number
[0124]
[0132] An offset that moves them away from the zero point can be applied to the quantization exponent. The amount of the offset ρ* can be adjusted through a validation set and used as a universal value for (e.g., all) videos, or it can be determined for each video / picture or transform block. The offset ρ* can be adaptable to the frequency position of the transformation coefficients in the transformed block. It can vary monotonically according to the quantization exponent i itself if a model other than a linear model is used for the rate modeling function.
[0125]
[0133] Figure 5 shows an exemplary encoder. As shown, a gradient transform 585 may be performed before inverse quantization. Figure 6 shows an exemplary decoder. As shown, a gradient transform 645 may be performed before inverse quantization. Since the inverse transform can be linear, in some examples the gradient transform may be performed after inverse quantization. The gradient transform may be applied before or after inverse quantization, and before the inverse transform.
[0126]
[0134] The gradient transform can be applied to the inversely quantized coefficients after inverse quantization. A rate prediction can be used where the gradient transform is equal to equation (7) and the set ρ*=0.04, because it can be universal. The flows from blocks 530 and 550 in Figure 5 can be seen as follows: x can be the transformed coefficients which are the output of block 525, and y=Q(x) can refer to block 530 as the quantization.
number
number
number
number
number
number
[0127]
[0135] Considering integer arithmetic, a=41 and N=10 can be used to achieve the effect of ρ*=0.04. This process may be the same for the decoder device in Figure 6.
[0128]
[0136] The gradient of the rate can be modeled by the reciprocal of the quantization exponent. An approximation of the gradient with respect to the rate can be as follows:
number
number
number
[0129]
[0137] When the gradient transform is applied to the inversely quantized coefficients after inverse quantization, the flows from blocks 530 and 550 in Figure 5 can be viewed as follows: x may be the transformed coefficients which are the output of block 525, and y=Q(x) may be the quantization pointing to block 530.
number
number
number
number
number
number
[0130]
[0138] In various examples, ρ* = 0.05 in Equation 8. Considering integer arithmetic, a = 51 and N = 10 are necessary to achieve the effect of ρ* = 0.05. This process may be the same as that for the decoder device in Figure 5.
[0131]
[0139] If the gradient transform is applied to the inversely quantized coefficients after inverse quantization, the shift selection in (9) may be as follows:
number
[0132]
[0140] (9) In y' i is the quantization exponent y i It is possible to show an auxiliary quantization exponent having an absolute value one exponential greater than y' i is y i It may have the same sign. The auxiliary quantization exponent can be calculated in (10) as follows: y' i =y i +(y i >0?1:-1) (10)
[0133]
[0141] In some cases, the shift parameter ρ* can be a value that enables the shift. ρ* can be a fixed value such as 42 or similar. N can be 10.
[0134]
[0142] In some examples, the shift parameter ρ* may be selected from multiple shift parameters. The shift may be selected based on one or more rules. The selection may be based on available information regarding the coefficients and / or their blocks and / or slices. If-else rules may be applied. In some examples, a three-level rule in (11) where the decision is based on absolute quantization values may be as follows: ρ*=ρ1, |y i |If a| ρ*=ρ², 0<|y i |≦a case (11) ρ*=0, y i If = 0
[0135]
[0143] The thresholds for a and / or the shift parameters ρ1, ρ2 may be predetermined and / or dynamically adaptable. For example, the thresholds for a and / or the shift parameters ρ1, ρ2 may be determined based on the training dataset.
[0136]
[0144] The shift parameter can be determined based on the quantization exponent. If the gradient transform is applied to the inversely quantized coefficients after inverse quantization, and the shift operation is chosen as the reciprocal proportional to the quantization exponent, then the example in equation (9) can be performed using the reciprocal proportional to the quantization exponent for ρ*. For example, the shift parameter ρ* in equation (12) can be chosen as follows:
number
[0137]
[0145] Equation (12) may use integer division during encoding and decoding. The division operation may be removed, and the lookup table
number
[0138]
[0146] For α = 63, the lookup table in Equation (14) can be implemented as follows: T = [0, 63, 31, 21, 15, 12, 10, 9, 7, 7, 6, 5, 5, 4, 4, 4, 3, 3, 3, 3, 3, 3, 2, 2, 2, 2, 2, 2, 2, 2, 1, …, 1] (14)
[0139]
[0147] The table for the shift parameter in Equation (13) can be changed using the training set. In various examples, the values of the table can be changed by the training set.
[0140]
[0148] To change the size of the shift parameter table, filtering can be applied as follows in (15): ρ* = T[|y i |], |y i | < a case (15) ρ* = 0, y i = 0 case
[0141]
[0149] The selection of the table size a can be an integer value between 0 and the table size (e.g., the maximum table size).
[0142]
[0150] The gradient of the rate can be modeled by the position of the coefficient to make the rate model recognize the frequency. The coefficients can come to the conversion block sequentially. The order (e.g., position) in which the coefficients come can indicate information about the frequency. For example, the first coefficient can represent the lowest frequency, and the second coefficient can represent a higher frequency. The order in which the coefficients come can be added to the model. As can be seen in (16), a conversion can be applied to the inverse-quantized coefficients.
Number
[0143]
[0151] The position of the coefficients can affect the gradient depending on the function. If the function is f(i)=1 / i, the flows from blocks 530 and 550 in Figure 5 can be as follows: y=Q(x) can be a quantization pointing to block 530,
number
number
number
number
number
number
[0144]
[0152] Considering integer arithmetic, a=164 and N=12 can be used to achieve the effect of ρ*=0.04. This process may be the same in the decoder device shown in Figure 6.
[0145]
[0153] In some cases, an optimized gradient transform may be used universally. In some cases, the optimality of the gradient transform may be adjusted per slice or per transform unit. The parameters of the gradient function may be adjusted to minimize the inverse quantization error. The parameters may be transmitted to the decoder. In some cases, the gradient function described herein may use ρ* for the transform unit. The parameters may be adjusted per slice or per video.
[0146]
[0154] In various examples, shifts (e.g., different selections of shifts) may be used. Shift selection may be applied (e.g., one by one) during the encoding stage, and shifts may be selected. The selected shifts may be signaled at the block, slice, and / or sequence levels.
[0147]
[0155] The variable x can be a transformation coefficient, which is the output of block 525. y=Q(x) can be a quantization pointing to block 530.
number
[0148]
[0156] a j ∈{0, a1, a2…, a k Regarding},
number
number
number
number
number
number
number
number
[0149]
[0157] In the decoder, before the gradient transformation, the selected gradient parameters can be read out and applied as follows: y can be the quantization exponent read from the bitstream,
number
number
[0150]
[0158]
number
number
number
number
number
[0151]
[0159] In some cases, the transformation may be applied after quantization (e.g., to the quantized exponents and / or inversely quantized coefficients) (e.g., without having their reciprocals before quantization). If a differentiable entropy (e.g., or rate) model is available, the transformation may be derived from the gradient of the entropy (e.g., or rate). The transformation may be derived by a manually designed function with available inputs and tuning the parameters with a validation set to minimize the quantization error. The transformation may be modeled using a neural network with desired available inputs, and the parameters may be tuned with a validation set to minimize the quantization error. The gradient transformation may be universal (e.g., one for all videos). The gradient transformation may be specific to a video / picture / transform block and may be indicated via one or more syntax elements in the video data (e.g., video bitstream).
[0152]
[0160] The system, method, and means are configured to shift the quantization center by a gradient of rates. A video coding device (e.g., a video encoding and / or video decoding device) may obtain a quantization exponent. The device may shift the quantization exponent based on a quantity that is the reciprocal of the quantization exponent. The device may apply the shifted quantization exponent to a conversion coefficient to generate an adjusted conversion coefficient.
[0153]
[0161] The device may determine that a quantity that is the reciprocal of the quantization exponent contains the gradient of the entropy rate. The device may determine that a quantity that is the reciprocal of the quantization exponent is based on a lookup table. The quantity may be a variable quantity, a fixed quantity, or a dynamically generated true gradient. The transformation coefficient may be a quantized transformation coefficient. The transformation coefficient may be an inversely quantized transformation coefficient. The quantization exponent may contain an inversely quantized exponent. The adjusted coefficient may contain an inversely quantized transformation coefficient.
[0154]
[0162] Systems, methods, and means may be used to shift the quantization center by a gradient of rates. A device (e.g., a video encoding device and / or video decoding device) may perform dequantization to obtain dequantized coefficients. A device may perform a gradient transform on the dequantized coefficients. A device may perform an inverse transform on the gradient-transformed dequantized coefficients.
[0155]
[0163] The device may determine a gradient transformation based on the entropy gradient. The device may determine a gradient transformation based on a function with parameters. The device may perform a gradient transformation on quantization exponents. The device may determine a transformation using a neural network with input parameters. The device may tune parameters using a validation set, which may minimize quantization errors. The device may perform a shift operation on inversely quantized coefficients, which may use shift parameters determined based on quantization exponents. The device may determine shift parameters using a lookup table based on the absolute value of the quantization exponents.
[0156]
[0164] The examples provided herein may assume that media content is streamed to a display device, but there is no particular limitation on the types of display devices that may benefit from the exemplary techniques described herein. For example, a display device may be a television, projector, mobile phone, tablet, etc. Furthermore, the exemplary techniques described herein may be applicable not only to streaming use cases but also to teleconferencing situations. In addition, the decoder and display as described herein may be separate devices or parts of the same device. For example, a set-top box may decode an incoming video stream and provide the decoded stream to a display device (for example, then) (for example, via HDMI), and information regarding viewing conditions such as viewing distance may be transmitted from the display device to the set-top box (for example, via HDMI).
[0157]
[0165] While features and elements are described above in specific combinations, those skilled in the art will understand that each feature or element can be used individually or in any combination with other features and elements. In addition, the methods described herein can be implemented in computer programs, software, or firmware embedded in computer-readable media for execution by a computer or processor. Examples of computer-readable media include electronic signals (transmitted via wired or wireless connections) and computer-readable storage media. Examples of computer-readable storage media include, but are not limited to, read-only memory (ROM), random access memory (RAM), registers, cache memory, semiconductor memory devices, magnetic media such as internal hard disks and removable disks, magneto-optical media, and optical media such as CD-ROM disks and digital multi-purpose disks (DVDs). A processor may be used, together with software, to implement a radio frequency transceiver for use with a WTRU, UE, terminal, base station, RNC, or any host computer.
Claims
1. A video encoding device, It is a processor, Obtaining the quantized exponent, Shifting the quantization exponent based on a quantity that is the reciprocal of the quantization exponent, To generate the adjusted conversion coefficients, the shifted quantization exponents are applied to the conversion coefficients, A video coding device equipped with a processor configured to perform the following.
2. The aforementioned processor, The device according to claim 1, wherein the quantity which is the reciprocal of the quantization exponent is determined to include the gradient of the entropy rate.
3. The aforementioned processor, The device according to claim 1, wherein the quantity which is the reciprocal of the quantization exponent is further configured to be determined based on a lookup table.
4. The device according to claim 1, wherein the quantity is one of a variable quantity, a fixed quantity, or a dynamically generated true gradient.
5. The device according to claim 1, wherein the conversion coefficient is a quantized conversion coefficient.
6. The device according to claim 1, wherein the conversion coefficient is an inversely quantized conversion coefficient.
7. The device according to claim 1, wherein the video encoding device is a video coding device or a video decoding device.
8. The device according to claim 1, wherein the quantization exponent includes an inverse quantization exponent, and the adjusted coefficient includes an inversely quantized transformation coefficient.
9. A method for a video coding device, wherein the method is It is a processor, Obtaining the quantized exponent, Shifting the quantization exponent based on a quantity that is the reciprocal of the quantization exponent, To generate the adjusted conversion coefficients, the shifted quantization exponents are applied to the conversion coefficients, A method, including a processor, configured to perform the following actions.
10. The method described above is The method according to claim 9, further comprising determining that the quantity which is the reciprocal of the quantization exponent includes the gradient of the entropy rate.
11. The method described above is The method according to claim 9, further comprising determining the quantity which is the reciprocal of the quantization exponent based on a lookup table.
12. The method according to claim 9, wherein the quantity is one of a variable quantity, a fixed quantity, or a dynamically generated true gradient.
13. The method according to claim 9, wherein the conversion coefficient is a quantized conversion coefficient.
14. The method according to claim 9, wherein the transformation coefficient is an inversely quantized transformation coefficient.
15. The method according to claim 9, wherein the quantization exponent includes an inverse quantization exponent, and the adjusted coefficient includes an inversely quantized transformation coefficient.