Ratio or code selection for LMCS

By selecting an appropriate ratio value and applying the gradient boosting decision tree algorithm in the video coding system, combined with machine learning models to optimize the coding process, the problem of low coding efficiency in existing technologies is solved, and more efficient video coding and decoding are achieved.

CN121970317APending Publication Date: 2026-05-01INTERDIGITAL CE PATENT HOLDINGS SAS
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
INTERDIGITAL CE PATENT HOLDINGS SAS
Filing Date
2024-09-20
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing video coding systems struggle to effectively select ratios or codewords in Luminance Mapping and Chromatography Scaling (LMCS), resulting in low coding efficiency.

Method used

By determining the encoding gain of candidate ratio values, selecting an appropriate ratio value, adjusting the number of codewords in the mapping function, and applying the gradient boosting decision tree algorithm to optimize the encoding process, combined with machine learning models to predict encoding efficiency, the codeword with the highest gain is selected for encoding.

Benefits of technology

It improves the efficiency and quality of video encoding, optimizes encoding gain selection, and enhances encoding efficiency and decoding accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121970317A_ABST
    Figure CN121970317A_ABST
Patent Text Reader

Abstract

Systems, methods, and instrumentalities may be configured for coding tree unit (CTU) grid shifting. An exemplary video decoding device may receive a picture encoded based on a ratio value. The device may receive a ratio value corresponding to the picture. The device may decode the picture based on the ratio value. The ratio value may correspond to a picture corresponding to a maximum gain. Decoding the picture based on the ratio value may include decoding the picture based on the maximum gain.
Need to check novelty before this filing date? Find Prior Art

Description

Cross-references to related applications on ratio or code selection for LMCS

[0001] This application claims the benefit of Provisional Patent Application No. 23306573.9, filed on September 21, 2023, the disclosure of which is incorporated herein by reference in its entirety. Background Technology

[0002] Video coding systems can be used to compress digital video signals, for example, to reduce the bandwidth required to store and / or transmit such signals. Video coding systems may include, for example, block-based, wavelet-based, and / or object-based systems. Summary of the Invention

[0003] The system, method, and instrumentality can be configured for ratio or code selection in Luminance Mapping and Chroma Scaling (LMCS). An exemplary video encoding apparatus can determine candidate ratio values ​​associated with a video image. The apparatus can determine the corresponding coding gain associated with each candidate ratio value. The apparatus can select a ratio value from the candidate ratio values ​​based on their respective associated coding gains. The apparatus can adjust the number of codewords in a mapping function associated with the video image based on the selected ratio value. The apparatus can encode the video image based on this mapping function.

[0004] This device can divide the range of luminance sample values ​​of a video image into luminance sample value segments, and these segments can correspond to corresponding sub-ranges of luminance values. The device can determine the number of codewords for each luminance sample value segment based on a selected ratio value. Based on the number of codewords determined for each luminance sample value segment, the device can map the luminance sample value segments to updated luminance sample values. The device can generate a mapping function based on the mapped luminance sample value segments. The device can include an indication of the selected ratio value in the video data, and the video image can be further encoded based on this indication.

[0005] The device can determine whether codewords are communicated via signals alone or whether these codewords are associated with predefined codes, and based on the determination that codewords are associated with predefined codes, it can include an indication of the selected ratio value in the video data.

[0006] The device can determine that a codeword is associated with a predefined code. The device can include an index associated with the codeword in the video data, and video images can be further encoded based on the codeword.

[0007] The device can determine at least one feature from mapped video image samples associated with a video image using at least one ratio value from candidate ratio values, and the coding gain can be determined based on the determined at least one feature. The at least one feature may include at least one of the following: a minimum value of a mapped video image sample, a maximum value of a mapped video image sample, a histogram of mapped video image samples segmented, a variance of mapped video image samples segmented, or the temporal variance between a mapped video image sample and a sample of a second video image different from the video image.

[0008] The device can apply the gradient boosting decision tree algorithm to candidate ratio values, and the encoding gain can be determined based on applying the gradient boosting decision tree algorithm to candidate ratio values.

[0009] An exemplary video decoding device can be configured to receive video images. The device can receive an indication of a ratio value corresponding to the video image. The device can adjust the number of codewords in a mapping function associated with the video image based on the ratio value. The device can then decode the video image based on the mapping function.

[0010] This device can divide the range of luminance sample values ​​of a video image into luminance sample value segments, and these segments can correspond to corresponding sub-ranges of luminance values. The device can determine the number of codewords for each luminance sample value segment based on a selected ratio. Based on the number of codewords determined for each luminance sample value segment, the device can map the luminance sample value segments to updated luminance sample values. The device can generate a mapping function based on the mapped luminance sample value segments.

[0011] This device can apply a mapping function to inter-frame predicted luminance samples. It can further decode video images based on this applied mapping function. The device can determine if a codeword belongs to a predefined codeword and can receive an indication of a ratio value based on this determination.

[0012] An exemplary video decoding device can receive images (e.g., video frames) encoded based on a ratio value. The device can receive the ratio value corresponding to the image. The device can decode the image based on the ratio value. The ratio value can correspond to an image with maximum gain. Decoding the image based on the ratio value can include decoding the image based on maximum gain.

[0013] An exemplary video decoding device can receive an image and an index. The device can determine a CODE (best codeword (CW)) from a predefined set of candidate CODEs based on the index. The device can then decode the image based on the CODE.

[0014] An exemplary video encoding apparatus may include receiving an image. The apparatus may estimate multiple encoding gains based on multiple ratio values. The apparatus may determine a maximum gain from these encoding gains. The apparatus may encode the image based on the maximum gain. Estimating the encoding gain may include testing ratio values ​​from the ratio values ​​on the image. The apparatus may determine that the ratio value corresponds to the maximum gain. Encoding the image based on the maximum gain may include encoding the image using the ratio value corresponding to the maximum gain.

[0015] An exemplary video encoding device can receive an image. The device can evaluate multiple candidate codes. The device can use a machine learning model to compute features to predict the encoding efficiency (e.g., gain) of the candidate codes. The device can compare the encoding efficiency of the candidate codes with the maximum gain. If the encoding efficiency of a candidate code is more efficient than the maximum gain, the device can select the code with the highest gain from the candidate codes. The device can select an index corresponding to the selected code. The device can encode the image based on this index. Attached Figure Description

[0016] Figure 1A is a system diagram illustrating an exemplary communication system that can implement one or more of the disclosed embodiments.

[0017] Figure 1B is a system diagram illustrating an exemplary wireless transmit / receive unit (WTRU) that can be used within the communication system illustrated in Figure 1A according to an embodiment.

[0018] Figure 1C is a system diagram illustrating an exemplary radio access network (RAN) and an exemplary core network (CN) that can be used within the communication system illustrated in Figure 1A according to an embodiment.

[0019] Figure 1D is a system diagram illustrating another exemplary RAN and another exemplary CN that can be used within the communication system illustrated in Figure 1A according to an embodiment.

[0020] Figure 2 illustrates an exemplary video encoder.

[0021] Figure 3 illustrates an exemplary video decoder.

[0022] Figure 4 illustrates a system example that can be implemented in various ways and with different features.

[0023] Figure 5 illustrates an exemplary flowchart for the inference used for ratio selection.

[0024] Figure 6 illustrates an exemplary flowchart of inference for a set of CODEs (e.g., the optimal set of segment codewords (CWs) for a video segment).

[0025] Figure 7 illustrates an exemplary flowchart of inference using features computed based on mapped video. Detailed Implementation

[0026] A more detailed understanding can be obtained from the following description, which is given by way of example in conjunction with the accompanying drawings.

[0027] Figure 1A is a diagram illustrating an exemplary communication system 100 that may implement one or more of the disclosed embodiments. The communication system 100 may be a multiple access system that provides content such as voice, data, video, messaging, and broadcasting to multiple wireless users. The communication system 100 enables multiple wireless users to access such content by sharing system resources, including wireless bandwidth. For example, the communication system 100 may employ one or more channel access methods, such as Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Frequency Division Multiple Access (FDMA), Orthogonal FDMA (OFDMA), Single Carrier FDMA (SC-FDMA), Zero Tail Unique Word DFT Extended OFDM (ZT UW DTS-s OFDM), Unique Word OFDM (UW-OFDM), Resource Block Filtered OFDM, Filter Bank Multicarrier (FBMC), etc.

[0028] As shown in Figure 1A, the communication system 100 may include wireless transmit / receive units (WTRUs) 102a, 102b, 102c, 102d, RAN 104 / 113, CN 106 / 115, public switched telephone network (PSTN) 108, Internet 110, and other networks 112. However, it should be understood that the disclosed embodiments contemplate any number of WTRUs, base stations, networks, and / or network elements. Each of the WTRUs 102a, 102b, 102c, 102d can be any type of device configured to operate and / or communicate in a wireless environment. For example, WTRUs 102a, 102b, 102c, and 102d (any of which may be referred to as a "station" and / or "STA") may be configured to transmit and / or receive wireless signals and may include user equipment (UE), mobile stations, fixed or mobile subscriber units, subscription-based units, pagers, cellular phones, personal digital assistants (PDAs), smartphones, laptops, netbooks, personal computers, wireless sensors, hotspots or Mi-Fi devices, Internet of Things (IoT) devices, watches or other wearable devices, head-mounted displays (HMDs), vehicles, drones, medical devices and applications (e.g., remote surgery), industrial devices and applications (e.g., robots and / or other wireless devices operating in industrial and / or automated processing chain environments), consumer electronics devices, devices operating on commercial and / or industrial wireless networks, etc. Any of WTRUs 102a, 102b, 102c, and 102d may be interchangeably referred to as a UE.

[0029] The communication system 100 may also include base station 114a and / or base station 114b. Each of base stations 114a and 114b may be any type of device configured to wirelessly interface with at least one of WTRUs 102a, 102b, 102c, and 102d to facilitate access to one or more communication networks (e.g., CN 106 / 115, Internet 110, and / or other networks 112). For example, base stations 114a and 114b may be base transceiver stations (BTS), Node-B, eNode B, home Node B, home eNode B, gNB, NR Node B, site controller, access point (AP), wireless router, etc. Although base stations 114a and 114b are each depicted as a single element, it should be understood that base stations 114a and 114b may include any number of interconnected base stations and / or network elements.

[0030] Base station 114a may be part of RAN 104 / 113, which may also include other base stations and / or network elements (not shown), such as base station controllers (BSCs), radio network controllers (RNCs), relay nodes, etc. Base station 114a and / or base station 114b may be configured to transmit and / or receive radio signals on one or more carrier frequencies (which may be referred to as cells (not shown)). These frequencies may be in licensed spectrum, unlicensed spectrum, or a combination of licensed and unlicensed spectrum. A cell may provide coverage of a specific geographic area (which may be relatively fixed or may vary over time). A cell may be further divided into cell sectors. For example, the cell associated with base station 114a may be divided into three sectors. Therefore, in one embodiment, base station 114a may include three transceivers, one for each sector of the cell. In embodiments, base station 114a may employ multiple-input multiple-output (MIMO) technology and may use multiple transceivers for each sector of the cell. For example, beamforming may be used to transmit and / or receive signals in a desired spatial direction.

[0031] Base stations 114a and 114b can communicate with one or more of WTRUs 102a, 102b, 102c, and 102d via air interface 116. Air interface 116 can be any suitable wireless communication link (e.g., radio frequency (RF), microwave, centimeter wave, micrometer wave, infrared (IR), ultraviolet (UV), visible light, etc.). Air interface 116 can be established using any suitable radio access technology (RAT).

[0032] More specifically, as described above, the communication system 100 can be a multiple access system and can employ one or more channel access schemes, such as CDMA, TDMA, FDMA, OFDMA, SC-FDMA, etc. For example, base stations 114a and WTRUs 102a, 102b, and 102c in RAN 104 / 113 can implement radio technologies such as Universal Mobile Telecommunications System (UMTS) Terrestrial Radio Access (UTRA), which can use Wideband CDMA (WCDMA) to establish air interfaces 115 / 116 / 117. WCDMA can include communication protocols such as High-Speed ​​Packet Access (HSPA) and / or Evolved HSPA (HSPA+). HSPA can include High-Speed ​​Downlink (DL) Packet Access (HSDPA) and / or High-Speed ​​UL Packet Access (HSUPA).

[0033] In the embodiment, base station 114a and WTRUs 102a, 102b, 102c can implement radio technologies such as Evolved UMTS Terrestrial Radio Access (E-UTRA), which can use Long Term Evolution (LTE) and / or LTE-Advanced (LTE-A) and / or LTE-Advanced Pro (LTE-A Pro) to establish air interface 116.

[0034] In the embodiment, base station 114a and WTRUs 102a, 102b, 102c may implement radio technologies such as NR radio access, which can use new radio (NR) to establish air interface 116.

[0035] In this embodiment, base station 114a and WTRUs 102a, 102b, and 102c can implement various radio access technologies. For example, base station 114a and WTRUs 102a, 102b, and 102c can implement both LTE and NR radio access together, for example, using the dual connectivity (DC) principle. Therefore, the air interface used by WTRUs 102a, 102b, and 102c can be characterized by various types of radio access technologies and / or transmissions to / from various types of base stations (e.g., eNBs and gNBs).

[0036] In other embodiments, base station 114a and WTRUs 102a, 102b, 102c may implement radio technologies such as IEEE 802.11 (i.e., Wi-Fi), IEEE 802.16 (i.e., WiMAX), CDMA2000, CDMA2000 1X, CDMA2000 EV-DO, Provisional Standard 2000 (IS-2000), Provisional Standard 95 (IS-95), Provisional Standard 856 (IS-856), Global System for Mobile Communications (GSM), Enhanced Data Rate GSM Evolution (EDGE), and GSM EDGE (GERAN).

[0037] Base station 114b in Figure 1A can be, for example, a wireless router, a home Node B, a home eNode B, or an access point, and can utilize any suitable RAT to facilitate wireless connectivity in localized areas (e.g., business premises, homes, vehicles, campuses, industrial facilities, air corridors (e.g., for drone use), roads, etc.). In one embodiment, base station 114b and WTRUs 102c, 102d can implement radio technologies such as IEEE 802.11 to establish a wireless local area network (WLAN). In another embodiment, base station 114b and WTRUs 102c, 102d can implement radio technologies such as IEEE 802.15 to establish a wireless personal area network (WPAN). In yet another embodiment, base station 114b and WTRUs 102c, 102d can utilize cellular-based RATs (e.g., WCDMA, CDMA2000, GSM, LTE, LTE-A, LTE-A Pro, NR, etc.) to establish picocells or femtocells. As shown in Figure 1A, base station 114b can have a direct connection to Internet 110. Therefore, base station 114b does not need to access Internet 110 via CN 106 / 115.

[0038] RAN 104 / 113 can communicate with CN 106 / 115, which can be any type of network configured to provide voice, data, application, and / or Voice over Internet Protocol (VoIP) services to one or more of WTRUs 102a, 102b, 102c, and 102d. Data can have different Quality of Service (QoS) requirements, such as different throughput requirements, latency requirements, fault tolerance requirements, reliability requirements, data throughput requirements, mobility requirements, etc. CN 106 / 115 can provide call control, billing services, location-based services, prepaid calling, internet connectivity, video distribution, etc., and / or perform advanced security functions such as user authentication. Although not shown in Figure 1A, it should be understood that RAN 104 / 113 and / or CN 106 / 115 can communicate directly or indirectly with other RANs using the same RAT as RAN 104 / 113 or a different RAT. For example, in addition to connecting to RAN 104 / 113, which can utilize NR radio technology, CN 106 / 115 can also communicate with another RAN (not shown) that uses GSM, UMTS, CDMA 2000, WiMAX, E-UTRA, or WiFi radio technology.

[0039] CN 106 / 115 may also act as a gateway for WTRU 102a, 102b, 102c, 102d to access PSTN 108, the Internet 110, and / or other networks 112. PSTN 108 may include a circuit-switched telephone network (POTS) providing conventional telephone service. The Internet 110 may include a global system of interconnected computer networks and devices using common communication protocols, such as Transmission Control Protocol (TCP), User Datagram Protocol (UDP), and / or Internet Protocol (IP) from the TCP / IP Internet Protocol suite. Network 112 may include wired and / or wireless communication networks owned and / or operated by other service providers. For example, network 112 may include another CN connected to one or more RANs using the same RAT as or a different RAT than RAN 104 / 113.

[0040] Some or all of the WTRUs 102a, 102b, 102c, and 102d in the communication system 100 may include multi-mode capabilities (e.g., WTRUs 102a, 102b, 102c, and 102d may include multiple transceivers for communicating with different wireless networks via different wireless links). For example, the WTRU 102c shown in Figure 1A may be configured to communicate with a base station 114a that may employ cellular-based radio technology and with a base station 114b that may employ IEEE 802 radio technology.

[0041] Figure 1B is a system diagram illustrating an exemplary WTRU 102. As shown in Figure 1B, the WTRU 102 may include a processor 118, a transceiver 120, a transmit / receive element 122, a speaker / microphone 124, a keyboard 126, a display / touchpad 128, non-removable memory 130, removable memory 132, a power supply 134, a Global Positioning System (GPS) chipset 136, and / or other peripheral devices 138, etc. It should be understood that, while remaining consistent with the embodiments, the WTRU 102 may include any sub-combination of the foregoing elements.

[0042] Processor 118 may be a general-purpose processor, a special-purpose processor, a conventional processor, a digital signal processor (DSP), multiple microprocessors, one or more microprocessors associated with a DSP core, a controller, a microcontroller, an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) circuit, any other type of integrated circuit (IC), a state machine, etc. Processor 118 may perform signal encoding, data processing, power control, input / output processing, and / or any other functions that enable WTRU 102 to operate in a wireless environment. Processor 118 may be coupled to transceiver 120, and transceiver 120 may be coupled to transmitting / receiving element 122. Although Figure 1B depicts processor 118 and transceiver 120 as separate components, it should be understood that processor 118 and transceiver 120 may be integrated together in an electronic package or on a chip.

[0043] Transmitting / receiving element 122 can be configured to transmit signals to or receive signals from a base station (e.g., base station 114a) via air interface 116. For example, in one embodiment, transmitting / receiving element 122 can be an antenna configured to transmit and / or receive RF signals. In embodiments, transmitting / receiving element 122 can be, for example, a transmitter / detector configured to transmit and / or receive IR, UV, or visible light signals. In yet another embodiment, transmitting / receiving element 122 can be configured to transmit and / or receive both RF and optical signals. It should be understood that transmitting / receiving element 122 can be configured to transmit and / or receive any combination of wireless signals.

[0044] Although the transmitting / receiving element 122 is depicted as a single element in FIG. 1B, the WTRU 102 may include any number of transmitting / receiving elements 122. More specifically, the WTRU 102 may employ MIMO technology. Thus, in one embodiment, the WTRU 102 may include two or more transmitting / receiving elements 122 (e.g., multiple antennas) for transmitting and receiving wireless signals via the air interface 116.

[0045] Transceiver 120 can be configured to modulate signals transmitted by transmitting / receiving element 122 and demodulate signals received by transmitting / receiving element 122. As described above, WTRU 102 can have multi-mode capability. Therefore, transceiver 120 can include multiple transceivers for enabling WTRU 102 to communicate via various RATs (e.g., NR and IEEE 802.11).

[0046] The processor 118 of WTRU 102 can be coupled to a speaker / microphone 124, a keyboard 126, and / or a display / touchpad 128 (e.g., a liquid crystal display (LCD) unit or an organic light-emitting diode (OLED) display unit) and can receive user input data therefrom. The processor 118 can also output user data to the speaker / microphone 124, keyboard 126, and / or display / touchpad 128. Furthermore, the processor 118 can access and store information from any type of suitable memory (e.g., non-removable memory 130 and / or removable memory 132). Non-removable memory 130 may include random access memory (RAM), read-only memory (ROM), a hard disk, or any other type of memory storage device. Removable memory 132 may include a subscriber identity module card (SIM), a memory stick, a secure digital storage (SD) card, etc. In other embodiments, the processor 118 can access and store information from memory that is not physically located on WTRU 102 (e.g., on a server or home computer (not shown)).

[0047] The processor 118 can receive power from the power supply 134 and can be configured to distribute and / or control power to other components in the WTRU 102. The power supply 134 can be any suitable device for powering the WTRU 102. For example, the power supply 134 may include one or more dry cell batteries (e.g., nickel-cadmium (NiCd), nickel-zinc (NiZn), nickel metal hydride (NiMH), lithium-ion (Li-ion), etc.), solar cells, fuel cells, etc.

[0048] The processor 118 may also be coupled to a GPS chipset 136, which may be configured to provide location information (e.g., longitude and latitude) about the current location of the WTRU 102. In addition to or instead of the information from the GPS chipset 136, the WTRU 102 may receive location information from base stations (e.g., base stations 114a, 114b) via air interface 116, and / or determine its location based on the timing of signals received from two or more nearby base stations. It should be understood that, while remaining consistent with the embodiments, the WTRU 102 may acquire location information using any suitable location determination method.

[0049] The processor 118 may also be coupled to other peripheral devices 138, which may include one or more software and / or hardware modules that provide additional features, functions, and / or wired or wireless connectivity. For example, peripheral devices 138 may include accelerometers, electronic compasses, satellite transceivers, digital cameras (for photos and / or videos), Universal Serial Bus (USB) ports, vibration devices, television transceivers, hands-free headsets, Bluetooth® modules, FM radio units, digital music players, media players, video game player modules, internet browsers, virtual reality and / or augmented reality (VR / AR) devices, activity trackers, etc. Peripheral devices 138 may include one or more sensors, such as gyroscopes, accelerometers, Hall effect sensors, magnetometers, orientation sensors, proximity sensors, temperature sensors, time sensors; geolocation sensors; altimeters, light sensors, touch sensors, magnetometers, barometers, gesture sensors, biometric sensors, and / or humidity sensors.

[0050] WTRU 102 may include a full-duplex radio for which the transmission and reception of some or all signals (e.g., associated with a specific subframe used for both UL (e.g., for transmission) and downlink (e.g., for reception)) may be concurrent and / or simultaneous. The full-duplex radio may include an interference management unit to reduce or substantially eliminate self-interference via hardware (e.g., a choke) or via signal processing by a processor (e.g., a separate processor (not shown) or via processor 118). In an embodiment, WTRU 102 may include a half-duplex radio for which the transmission and reception of some or all signals (e.g., associated with a specific subframe used for UL (e.g., for transmission) or downlink (e.g., for reception)) may be concurrent and / or simultaneous.

[0051] Figure 1C is a system diagram illustrating RAN 104 and CN 106 according to an embodiment. As described above, RAN 104 can communicate with WTRUs 102a, 102b, and 102c via air interface 116 using E-UTRA radio technology. RAN 104 can also communicate with CN 106.

[0052] RAN 104 may include eNode-Bs 160a, 160b, and 160c; however, it should be understood that RAN 104 may include any number of eNode-Bs while remaining consistent with the embodiments. eNode-Bs 160a, 160b, and 160c may each include one or more transceivers for communicating with WTRUs 102a, 102b, and 102c via air interface 116. In one embodiment, eNode-Bs 160a, 160b, and 160c may implement MIMO technology. Therefore, for example, eNode-B 160a may use multiple antennas to transmit radio signals to and / or receive radio signals from WTRU 102a.

[0053] Each of the eNode-B 160a, 160b, and 160c can be associated with a specific cell (not shown) and can be configured to handle radio resource management decisions, handover decisions, user scheduling in UL and / or DL, etc. As shown in Figure 1C, the eNode-B 160a, 160b, and 160c can communicate with each other via the X2 interface.

[0054] The CN 106 shown in Figure 1C may include a Mobility Management Entity (MME) 162, a Serving Gateway (SGW) 164, and a Packet Data Network (PDN) Gateway (or PGW) 166. While each of the foregoing elements is depicted as part of CN 106, it should be understood that any of these elements may be owned and / or operated by an entity other than the CN operator.

[0055] The MME 162 can connect to each of the eNode-Bs 162a, 162b, and 162c in RAN 104 via the S1 interface and can act as a control node. For example, the MME 162 can be responsible for authenticating users of WTRUs 102a, 102b, and 102c, bearer activation / deactivation, and selecting a specific service gateway during the initial attachment of WTRUs 102a, 102b, and 102c. The MME 162 can provide control plane functions for handover between RAN 104 and other RANs (not shown) employing other radio technologies (such as GSM and / or WCDMA).

[0056] The SGW 164 can connect to each of the eNode Bs 160a, 160b, and 160c in RAN 104 via the S1 interface. The SGW 164 typically routes and forwards user data packets to / from WTRUs 102a, 102b, and 102c. The SGW 164 can perform other functions, such as anchoring the user plane during inter-eNode B handover, triggering paging when DL data is available for WTRUs 102a, 102b, and 102c, and managing and storing the context of WTRUs 102a, 102b, and 102c.

[0057] SGW 164 can connect to PGW 166, which can provide WTRU 102a, 102b, and 102c with access to packet-switched networks (such as Internet 110) to facilitate communication between WTRU 102a, 102b, and 102c and IP-enabled devices.

[0058] CN 106 can facilitate communication with other networks. For example, CN 106 can provide WTRUs 102a, 102b, and 102c with access to a circuit-switched network (e.g., PSTN 108) to facilitate communication between WTRUs 102a, 102b, and 102c and traditional landline communication equipment. For example, CN 106 may include an IP gateway (e.g., an IP Multimedia Subsystem (IMS) server) that acts as an interface between CN 106 and PSTN 108, or can communicate with it. Furthermore, CN 106 can provide WTRUs 102a, 102b, and 102c with access to other networks 112, which may include other wired and / or wireless networks owned and / or operated by other service providers.

[0059] Although the WTRU is described as a wireless terminal in Figures 1A to 1D, it is conceivable that in some representative embodiments such a terminal may (e.g., temporarily or permanently) use a wired communication interface with a communication network.

[0060] In a representative embodiment, the other network 112 may be a WLAN.

[0061] In Infrastructure Basic Services Set (BSS) mode, a WLAN may have an access point (AP) for the BSS and one or more stations (STAs) associated with that AP. The AP may have access or an interface to a distribution system (DS) or another type of wired / wireless network that transmits traffic into and / or out of the BSS. Traffic flows originating outside the BSS destined for a STA can reach the AP and be delivered to the STA. Traffic flows originating from a STA destined for a destination outside the BSS can be sent to the AP for delivery to the appropriate destination. Traffic flows between STAs within the BSS can be transmitted, for example, via the AP, where a source STA can send a traffic flow to the AP, and the AP can deliver the traffic flow to the destination STA. Traffic flows between STAs within the BSS can be considered and / or referred to as peer-to-peer traffic flows. Peer-to-peer traffic flows can be transmitted between source and destination STAs (e.g., directly between them) using Direct Link Establishment (DLS). In some representative embodiments, the DLS may use 802.11e DLS or 802.11z Tunneled DLS (TDLS). A WLAN using the Standalone BSS (IBSS) mode may not have an access point (AP), and STAs within the IBSS or using the IBSS (e.g., all STAs) can communicate directly with each other. The IBSS communication mode may sometimes be referred to as an "ad-hoc" communication mode in this article.

[0062] When operating in 802.11ac infrastructure mode or a similar mode, the AP can transmit beacons on a fixed channel (e.g., the primary channel). The primary channel can have a fixed width (e.g., a 20 MHz wide bandwidth) or a width dynamically set via signaling. The primary channel can be the operating channel of the BSS and can be used by STAs to establish connections with the AP. In some representative embodiments, such as in an 802.11 system, Carrier Sense Multiple Access with Collision Avoidance (CSMA / CA) can be implemented. For CSMA / CA, STAs including the AP (e.g., each STA) can listen on the primary channel. If a particular STA detects / investigates and / or determines that the primary channel is busy, that particular STA can back off. In a given BSS, there can be only one STA (e.g., only one station) transmitting at any given time.

[0063] High-throughput (HT) STAs can communicate using a 40 MHz wide channel, for example, by combining a primary 20 MHz channel with adjacent or non-adjacent 20 MHz channels to form a 40 MHz wide channel.

[0064] Ultra-high throughput (VHT) STAs can support channels with widths of 20MHz, 40MHz, 80MHz, and / or 160MHz. 40MHz and / or 80MHz channels can be formed by combining consecutive 20MHz channels. A 160MHz channel can be formed by combining eight consecutive 20MHz channels, or by combining two non-consecutive 80MHz channels (which can be referred to as an 80+80 configuration). For the 80+80 configuration, data can be delivered after channel coding through a segment parser, which splits the data into two streams. Inverse Fast Fourier Transform (IFFT) processing and time-domain processing can be performed separately on each stream. These streams can be mapped onto two 80MHz channels, and the data can be transmitted by the transmitting STA. At the receiver of the receiving STA, the above operations for the 80+80 configuration can be reversed, and the combined data can be sent to the Media Access Control (MAC).

[0065] Sub-1 GHz operating modes are supported by 802.11af and 802.11ah. Compared to those used in 802.11n and 802.11ac, the channel operating bandwidth and carrier in 802.11af and 802.11ah are reduced. 802.11af supports 5 MHz, 10 MHz, and 20 MHz bandwidths in the TV white space (TVWS) spectrum, while 802.11ah supports 1 MHz, 2 MHz, 4 MHz, 8 MHz, and 16 MHz bandwidths using non-TVWS spectrum. According to representative embodiments, 802.11ah can support instrument-type control / machine-type communication, such as MTC devices in macro coverage areas. MTC devices may have certain capabilities, such as limited capabilities, including supporting (e.g., only supporting) certain and / or limited bandwidths. MTC devices may include batteries with a battery life exceeding a threshold (e.g., to maintain a very long battery life).

[0066] WLAN systems supporting multiple channels and channel bandwidths (e.g., 802.11n, 802.11ac, 802.11af, and 802.11ah) include channels that can be designated as the primary channel. The primary channel can have a bandwidth equal to the maximum common operating bandwidth supported by all STAs in the BSS. The bandwidth of the primary channel can be set and / or limited by the STAs operating in the BSS that support the minimum bandwidth operating mode. In the 802.11ah example, for STAs supporting (e.g., only supporting) the 1 MHz mode (e.g., MTC type devices), the primary channel can be 1 MHz wide, even if the AP and other STAs in the BSS support 2 MHz, 4 MHz, 8 MHz, 16 MHz, and / or other channel bandwidth operating modes. Carrier Sense and / or Network Assignment Vector (NAV) settings can depend on the status of the primary channel. If the primary channel is busy, for example, because an STA (supporting only the 1 MHz operating mode) is transmitting to the AP, the entire available band can be considered busy, even if most of the band remains idle and may be available.

[0067] In the United States, the available frequency band for 802.11ah is 902 MHz to 928 MHz. In South Korea, the available frequency band is 917.5 MHz to 923.5 MHz. In Japan, the available frequency band is 916.5 MHz to 927.5 MHz. The total available bandwidth for 802.11ah varies from 6 MHz to 26 MHz depending on the country code.

[0068] Figure 1D is a system diagram illustrating RAN 113 and CN 115 according to an embodiment. As described above, RAN 113 can communicate with WTRUs 102a, 102b, and 102c via air interface 116 using NR radio technology. RAN 113 can also communicate with CN 115.

[0069] RAN 113 may include gNBs 180a, 180b, and 180c; however, it should be understood that RAN 113 may include any number of gNBs while remaining consistent with the embodiments. gNBs 180a, 180b, and 180c may each include one or more transceivers for communicating with WTRUs 102a, 102b, and 102c via air interface 116. In one embodiment, gNBs 180a, 180b, and 180c may implement MIMO technology. For example, gNBs 180a and 180b may utilize beamforming to transmit signals to and / or receive signals from gNBs 180a, 180b, and 180c. Thus, for example, gNB 180a may use multiple antennas to transmit radio signals to and / or receive radio signals from WTRU 102a. In an embodiment, gNBs 180a, 180b, and 180c may implement carrier aggregation technology. For example, gNB 180a can transmit multiple component carriers (not shown) to WTRU 102a. A subset of these component carriers may be on unlicensed spectrum, while the remaining component carriers may be on licensed spectrum. In embodiments, gNBs 180a, 180b, and 180c may implement Coordinated Multipoint (CoMP) technology. For example, WTRU 102a may receive coordinated transmissions from gNBs 180a and 180b (and / or gNB 180c).

[0070] WTRUs 102a, 102b, and 102c can communicate with gNBs 180a, 180b, and 180c using transmissions associated with scalable parameter sets. For example, OFDM symbol spacing and / or OFDM subcarrier spacing can vary depending on different transmissions, different cells, and / or different portions of the radio transmission spectrum. WTRUs 102a, 102b, and 102c can communicate with gNBs 180a, 180b, and 180c using subframes or transmission time intervals (TTIs) of various or scalable lengths (e.g., containing different numbers of OFDM symbols and / or varying absolute time lengths).

[0071] gNBs 180a, 180b, and 180c can be configured to communicate with WTRUs 102a, 102b, and 102c in standalone and / or non-standalone configurations. In standalone configuration, WTRUs 102a, 102b, and 102c can communicate with gNBs 180a, 180b, and 180c without simultaneously accessing other RANs (e.g., eNode-Bs 160a, 160b, and 160c). In standalone configuration, WTRUs 102a, 102b, and 102c can use one or more of gNBs 180a, 180b, and 180c as mobility anchors. In standalone configuration, WTRUs 102a, 102b, and 102c can communicate with gNBs 180a, 180b, and 180c using signals in unlicensed frequency bands. In a non-standalone configuration, WTRUs 102a, 102b, and 102c can communicate / connect with gNBs 180a, 180b, and 180c while simultaneously communicating / connecting with another RAN (e.g., eNode-Bs 160a, 160b, and 160c). For example, WTRUs 102a, 102b, and 102c can implement DC principles to communicate substantially simultaneously with one or more gNBs 180a, 180b, and 180c, as well as one or more eNode-Bs 160a, 160b, and 160c. In a non-standalone configuration, eNode-Bs 160a, 160b, and 160c can act as mobility anchors for WTRUs 102a, 102b, and 102c, and gNBs 180a, 180b, and 180c can provide additional coverage and / or throughput to serve WTRUs 102a, 102b, and 102c.

[0072] Each of gNBs 180a, 180b, and 180c can be associated with a specific cell (not shown) and can be configured to handle radio resource management decisions, handover decisions, user scheduling in UL and / or DL, network slicing support, dual connectivity, interoperability between NR and E-UTRA, routing of user plane data to User Plane Functions (UPF) 184a and 184b, and routing of control plane information to Access and Mobility Management Functions (AMF) 182a and 182b, etc. As shown in Figure 1D, gNBs 180a, 180b, and 180c can communicate with each other via the Xn interface.

[0073] The CN 115 shown in Figure 1D may include at least one AMF 182a, 182b, at least one UPF 184a, 184b, at least one Session Management Function (SMF) 183a, 183b, and possibly a Data Network (DN) 185a, 185b. Although each of the foregoing elements is depicted as part of the CN 115, it should be understood that any of these elements may be owned and / or operated by an entity other than the CN operator.

[0074] AMF 182a and 182b can connect to one or more gNBs 180a, 180b, and 180c in RAN 113 via the N2 interface and can act as control nodes. For example, AMF 182a and 182b can be responsible for authenticating users of WTRU 102a, 102b, and 102c, supporting network slicing (e.g., handling different PDU sessions with different requirements), selecting specific SMF 183a and 183b, managing registration areas, terminating NAS signaling, mobility management, etc. Network slicing can be used by AMF 182a and 182b to customize CN support for WTRU 102a, 102b, and 102c based on the service types being used by WTRU 102a, 102b, and 102c. For example, different network slices can be established for different use cases, such as services relying on Ultra Reliable Low Latency (URLLC) access, services relying on Enhanced Massive Mobile Broadband (eMBB) access, and services for Machine Type Communication (MTC) access. AMF 162 can provide control plane functionality for handover between RAN 113 and other RANs (not shown) that employ other radio technologies, such as LTE, LTE-A, LTE-A Pro and / or non-3GPP access technologies, such as WiFi.

[0075] SMFs 183a and 183b can connect to AMFs 182a and 182b in CN 115 via the N11 interface. SMFs 183a and 183b can also connect to UPFs 184a and 184b in CN 115 via the N4 interface. SMFs 183a and 183b can select and control UPFs 184a and 184b, and configure traffic routing through UPFs 184a and 184b. SMFs 183a and 183b can perform other functions, such as managing and allocating UE IP addresses, managing PDU sessions, controlling policy enforcement and QoS, and providing downlink data notifications. PDU session types can be IP-based, non-IP-based, Ethernet-based, etc.

[0076] UPF 184a and 184b can connect to one or more gNBs 180a, 180b, and 180c in RAN 113 via the N3 interface. They can provide WTRU 102a, 102b, and 102c with access to packet-switched networks (such as Internet 110) to facilitate communication between WTRU 102a, 102b, and 102c and IP-enabled devices. UPF 184 and 184b can perform other functions such as routing and forwarding packets, enforcing user plane policies, supporting multi-homed PDU sessions, handling user plane QoS, buffering downlink packets, and providing mobility anchoring.

[0077] CN 115 can facilitate communication with other networks. For example, CN 115 may include or be able to communicate with an IP gateway (e.g., an IP Multimedia Subsystem (IMS) server) that acts as an interface between CN 115 and PSTN 108. Furthermore, CN 115 can provide WTRUs 102a, 102b, and 102c with access to other networks 112, which may include other wired and / or wireless networks owned and / or operated by other service providers. In one embodiment, WTRUs 102a, 102b, and 102c can be connected to DNs 185a and 185b via UPFs 184a and 184b through their N3 interfaces and the N6 interface between UPFs 184a and 184b and local data networks (DNs) 185a and 185b.

[0078] Based on Figures 1A to 1D and the corresponding descriptions of Figures 1A to 1D, one or more of the functions described below can be performed by one or more emulation devices (not shown): WTRU 102a to 102d, base stations 114a and 114b, eNode-B 160a and 160c, MME 162, SGW 164, PGW 166, gNB 180a to 180c, AMF 182a and 182b, UPF 184a and 184b, SMF 183a and 183b, DN 185a and 185b, and / or any other devices described herein. An emulation device can be one or more devices configured to emulate one or more of the functions described herein. For example, an emulation device can be used to test other devices and / or simulate network and / or WTRU functions.

[0079] Simulation devices can be designed to perform one or more tests on other devices in laboratory and / or carrier network environments. For example, one or more simulation devices can perform one or more functions when implemented and / or deployed, wholly or partially, as part of a wired and / or wireless communication network to test other devices within the communication network. One or more simulation devices can perform one or more functions when temporarily implemented / deployed as part of a wired and / or wireless communication network. Simulation devices can be directly coupled to another device for testing purposes and / or for performing tests using over-the-air wireless communication.

[0080] One or more simulation devices may perform one or more functions without being implemented / deployed as part of a wired and / or wireless communication network. For example, simulation devices may be used to perform tests on one or more components in a test scenario in a test laboratory and / or in a non-deployed (e.g., testing) wired and / or wireless communication network. One or more simulation devices may be test devices. Simulation devices may transmit and / or receive data using direct RF coupling and / or wireless communication via an RF circuit system (e.g., which may include one or more antennas).

[0081] This application describes various aspects, including tools, features, examples, models, methods, etc. Many of these aspects are described in detail, and often in a manner that may sound restrictive, at least to demonstrate individual characteristics. However, this is for the purpose of clarity and does not limit the application or scope of these aspects. In fact, all the different aspects can be combined and interchanged to provide further aspects. Furthermore, these aspects can be combined and interchanged with aspects described in prior applications.

[0082] The aspects described and contemplated in this application can be implemented in many different forms. Figures 5 through 11 described herein provide some examples, but other examples are also contemplated. The discussion of Figures 5 through 11 does not limit the breadth of implementations. At least one aspect generally relates to video encoding and decoding, and at least another aspect generally relates to transmitting a generated or encoded bitstream. These and other aspects can be implemented as methods, apparatus, computer-readable storage media having instructions thereon stored thereon for encoding or decoding video data according to any of the described methods, and / or computer-readable storage media having a bitstream generated according to any of the described methods stored thereon.

[0083] In this application, the terms "reconstruction" and "decoding" are used interchangeably, the terms "pixel" and "sample" are used interchangeably, and the terms "image," "picture," and "frame" are used interchangeably.

[0084] This document describes various methods, and each method includes one or more steps or actions for implementing the described method. Unless a specific order of steps or actions is required for the correct operation of the method, the order and / or use of specific steps and / or actions can be modified or combined. Furthermore, terms such as "first" and "second" can be used in various examples to modify elements, components, steps, operations, etc., such as "first decoding" and "second decoding." Unless specifically required, the use of such terms does not imply the order of the modified operations. Therefore, in this example, the first decoding need not be performed before the second decoding, and can occur, for example, before, during, or within a time period overlapping with the second decoding.

[0085] The various methods and other aspects described in this application can be used to modify modules, such as the decoding modules of the video encoder 200 and decoder 300 shown in Figures 2 and 3. Furthermore, the subject matter disclosed herein can be applied to, for example, any type, format, or version of video coding, whether described in a standard or recommendation, whether pre-existing or future-developed, and any extensions to such standards and recommendations. Unless otherwise stated or technically excluded, the aspects described in this application can be used individually or in combination.

[0086] The examples described in this application use various numerical values, such as 1, 2, 4, 7, 8, 16, 32, 64, etc. These and other specific values ​​are for illustrative purposes, and the aspects described are not limited to these specific values.

[0087] Figure 2 is a diagram illustrating an exemplary video encoder 200. Variations of the exemplary encoder 200 are contemplated, but for clarity, the encoder 200 is described below without describing all anticipated variations.

[0088] Prior to encoding, the video sequence may undergo pre-processing (201), such as applying color transformations to the input color image (e.g., a conversion from RGB 4:4:4 to YCbCr 4:2:0), or performing remapping of the input image components to obtain a more resilient signal distribution to compression (e.g., using histogram equalization on one of the color components). Metadata (e.g., which may include film grain parameters determined through preprocessing as described herein) may be associated with preprocessing and appended to the bitstream.

[0089] In encoder 200, the image is encoded by encoder elements as described below. The image to be encoded is divided (202) and processed, for example, in units of coding units (Cu). Each unit is encoded, for example, using an intra-frame mode or an inter-frame mode. When a unit is encoded in intra-frame mode, it performs intra-frame prediction (260). In inter-frame mode, motion estimation (275) and compensation (270) are performed. The encoder determines (205) which of the intra-frame or inter-frame modes to use to encode the unit, and indicates the intra / inter-frame decision, for example, by a prediction mode flag. The prediction residual is calculated, for example, by subtracting (210) the prediction block from the original image block.

[0090] The predicted residual is then transformed (225) and quantized (230). The quantized transform coefficients, along with the motion vector and other syntax elements, are entropy encoded (245) to output a bitstream. The encoder can skip the transform and apply quantization directly to the untransformed residual signal. The encoder can bypass both the transform and quantization, i.e., the residual is directly encoded without applying either the transform or quantization process.

[0091] The encoder decodes the encoded blocks to provide a reference for further prediction. The quantized transform coefficients are dequantized (240) and inverse transformed (250) to decode the prediction residuals. The decoded prediction residuals and the predicted blocks are combined (255) to reconstruct the image blocks. A loop filter (265) is applied to the reconstructed image to perform, for example, deblocking / SAO (Sample Adaptive Shift) filtering to reduce coding artifacts. The filtered image is stored in a reference image buffer (280).

[0092] Figure 3 is a diagram illustrating an example of a video decoder. In the exemplary decoder 300, the bitstream is decoded by the decoder elements as described below. The video decoder 300 typically performs a decoding process that is the inverse of the encoding process described in Figure 2. The encoder 200 also typically performs video decoding as part of the video data encoding.

[0093] Specifically, the input to the decoder includes a video bitstream, which can be generated by the video encoder 200. First, the bitstream is entropy-decoded (330) to obtain transform coefficients, motion vectors, and other encoded information. Image segmentation information indicates how the image is segmented. Therefore, the decoder can segment (335) the image based on the decoded image segmentation information. The transform coefficients are dequantized (340) and inverse transformed (350) to decode the prediction residuals. The decoded prediction residuals and predicted blocks are combined (355) to reconstruct image blocks. The predicted blocks can be obtained from intra-frame prediction (360) or motion-compensated prediction (i.e., inter-frame prediction) (375) (370). A loop filter (365) is applied to the reconstructed image. The filtered image is stored in a reference image buffer (380).

[0094] The decoded image can be further processed by post-decoding (385), such as inverse color transformation (e.g., conversion from YcbCr4:2:0 to RGB 4:4:4) or inverse remapping of the process that performs the inverse of the remapping process performed in pre-encoding (201). Post-decoding can use metadata exported in pre-encoding and signaled in the bitstream. In the example, the decoded image (e.g., after applying a loop filter (365) and / or after post-decoding (385) in the case of post-decoding) can be sent to a display device for presentation to the user.

[0095] Figure 4 is a diagram illustrating system examples that can implement the various aspects and examples described herein. System 400 can be implemented as a device including the various components described below and configured to perform one or more aspects described in this document. Examples of such devices include, but are not limited to, various electronic devices such as personal computers, laptop computers, smartphones, tablet computers, digital multimedia set-top boxes, digital television receivers, personal video recording systems, networked home appliances, and servers. Elements of system 400 can be implemented individually or in combination in a single integrated circuit (IC), multiple ICs, and / or discrete devices. For example, in at least one example, the processing and encoder / decoder elements of system 400 are distributed across multiple ICs and / or discrete devices. In many examples, system 400 is communicatively coupled to one or more other systems or other electronic devices, for example, via a communication bus or through dedicated input and / or output ports. In many examples, system 400 is configured to implement one or more aspects described in this document.

[0096] System 400 includes at least one processor 410 configured to execute instructions loaded thereon to implement various aspects, such as those described in this document. Processor 410 may include embedded memory, input / output interfaces, and a variety of other circuit systems known in the art. System 400 includes at least one memory 420 (e.g., a volatile memory device and / or a non-volatile memory device). System 400 includes a storage device 440, which may include non-volatile memory and / or volatile memory, including but not limited to electrically erasable programmable read-only memory (EEPROM), read-only memory (ROM), programmable read-only memory (PROM), random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), flash memory, disk drives, and / or optical disk drives. Storage device 440 may include internal storage devices, attached storage devices (including removable and non-removable storage devices), and / or network-accessible storage devices, as non-limiting examples.

[0097] System 400 includes an encoder / decoder module 430, which is configured, for example, to process data to provide encoded or decoded video, and the encoder / decoder module 430 may include its own processor and memory. The encoder / decoder module 430 represents a module that can be included in a device to perform encoding and / or decoding functions. It is well known that a device may include one or both encoding and decoding modules. Furthermore, the encoder / decoder module 430 may be implemented as a separate element of system 400, or it may be incorporated within processor 410 as a combination of hardware and software, as is known to those skilled in the art.

[0098] Program code loaded onto processor 410 or encoder / decoder 430 to execute the various aspects described in this document may be stored in storage device 440 and subsequently loaded into memory 420 for execution by processor 410. According to various examples, one or more of processor 410, memory 420, storage device 440, and encoder / decoder module 430 may store one or more items of various kinds during the execution of the processes described in this document. Such stored items may include, but are not limited to, input video, decoded video or portions thereof, bitstreams, matrices, variables, and intermediate or final results from processing equations, formulas, operations, and operational logic.

[0099] In some examples, the memory within processor 410 and / or encoder / decoder module 430 is used to store instructions and provide working memory for the processing required during encoding or decoding. In other examples, however, external memory (e.g., the processing device may be processor 410 or encoder / decoder module 430) is used for one or more of these functions. External memory may be, for example, memory 420 and / or storage device 440, such as volatile memory and / or non-volatile flash memory. In several examples, external non-volatile flash memory is used to store, for example, the operating system of a television. In at least one example, fast external volatile memory such as RAM is used as working memory for video encoding and decoding operations.

[0100] Inputs to the components of system 400 can be provided via a variety of input devices as indicated in box 445. Such input devices include, but are not limited to, (i) an RF section that receives radio frequency (RF) signals, for example, transmitted over the air by a broadcaster, (ii) component input terminals (or sets of COMP input terminals), (iii) universal serial bus (USB) input terminals, and / or (iv) high-definition multimedia interface (HDMI) input terminals. Other examples not shown in Figure 4 include composite video.

[0101] In various examples, the input device of block 445 has corresponding associated input processing elements known in the art. For example, the RF section may be associated with elements suitable for: (i) selecting a desired frequency (also known as selecting a signal, or band-limiting a signal to a frequency band), (ii) down-converting the selected signal, (iii) band-limiting it again to a narrower frequency band to select, for example, a signal band (which may be referred to as a channel in some examples), (iv) demodulating the down-converted and band-limited signal, (v) performing error correction, and / or (vi) demultiplexing to select the desired data packet stream. The RF section in various examples includes one or more elements performing these functions, such as frequency selectors, signal selectors, band limiters, channel selectors, filters, downconverters, demodulators, error correctors, and demultiplexers. The RF section may include a tuner that performs multiple of these functions, such as down-converting a received signal to a lower frequency (e.g., intermediate frequency or near-baseband frequency) or baseband. In one set-top box example, the RF section and its associated input processing elements receive RF signals transmitted via a wired (e.g., cable) medium and perform frequency selection by filtering, down-converting, and re-filtering to the desired frequency band. Various examples rearrange the order of the above (and other) components, remove some of these components, and / or add other components that perform similar or different functions. Adding components may include inserting components between existing components, such as inserting amplifiers and analog-to-digital converters. In many examples, the RF section includes an antenna.

[0102] USB and / or HDMI terminals may include corresponding interface processors for connecting system 400 to other electronic devices across USB and / or HDMI connections. It should be understood that various aspects of input processing (e.g., Reed-Solomon error correction) may be implemented, for example, within a separate input processing IC or within processor 410 as needed. Similarly, aspects of USB or HDMI interface processing may be implemented, for example, within a separate interface IC or within processor 410 as needed. The demodulated, error-corrected, and demultiplexed stream is provided as needed to various processing elements, such as processor 410 and an encoder / decoder 430 operating in combination with memory and storage elements, to process the data stream for presentation on an output device.

[0103] Various components of system 400 can be housed within an integrated housing. Within the integrated housing, various components can be interconnected and transmit data therebetween using suitable connection means 425 (e.g., internal buses known in the art, including inter-IC (I2C) buses, wiring, and printed circuit boards).

[0104] System 400 includes a communication interface 450 that enables communication with other devices via a communication channel 460. The communication interface 450 may include, but is not limited to, a transceiver configured to send and receive data via the communication channel 460. The communication interface 450 may include, but is not limited to, a modem or network interface card (NIC), and the communication channel 460 may be implemented, for example, within a wired and / or wireless medium.

[0105] In several examples, data is provided to system 400 via streaming or otherwise using a wireless network such as Wi-Fi (e.g., IEEE 802.11). In these examples, the Wi-Fi signal is received via a communication channel 460 and a communication interface 450 adapted for Wi-Fi communication. The communication channel 460 in these examples is typically connected to an access point or router that provides access to external networks, including the Internet, to allow streaming applications and other over-the-top services to communicate. Other examples use a set-top box to deliver streaming data to the system 400 via an HDMI connection in input box 445. Still other examples use an RF connection in input box 445 to provide streaming data to the system 400. As mentioned above, several examples provide data in a non-streaming manner. Furthermore, several examples use wireless networks other than Wi-Fi, such as cellular networks or Bluetooth® networks.

[0106] System 400 can provide output signals to a variety of output devices, including a display 475, a speaker 485, and other peripheral devices 495. Various examples of the display 475 include one or more of a touchscreen display, an organic light-emitting diode (OLED) display, a flexible display, and / or a foldable display. The display 475 can be used in televisions, tablets, laptops, cellular phones (mobile phones), or other devices. The display 475 can also be integrated with other components (e.g., as in a smartphone) or separate (e.g., as in an external monitor for a laptop). Other peripheral devices 495 include, in various examples, one or more of a standalone digital video disc (or digital multifunction disc) (DVD, for both terms), an optical disc player, a stereo system, and / or a lighting system. Various examples utilize one or more peripheral devices 495 that provide functionality based on the output of system 400. For example, an optical disc player performs the function of playing the output of system 400.

[0107] In various examples, signaling such as AV.Link, Consumer Electronics Control (CEC), or other communication protocols is used to transmit control signals between system 400 and display 475, speaker 485, or other peripheral devices 495. These communication protocols enable device-to-device control with or without user intervention. Output devices can be communicatively coupled to system 400 via dedicated connections through corresponding interfaces 470, 480, and 490. Alternatively, output devices can be connected to system 400 via communication interface 450 using communication channel 460. Display 475 and speaker 485 can be integrated into a single unit within an electronic device such as a television, along with other components of system 400. In various examples, display interface 470 includes a display driver, such as a timing controller (TCon) chip.

[0108] Display 475 and speaker 485 may alternatively be separate from one or more other components, for example, if the RF section of input 445 is part of a separate set-top box. In various examples where display 475 and speaker 485 are external components, the output signal may be provided via a dedicated output connection, such as an HDMI port, USB port, or COMP output.

[0109] The example can be implemented by computer software (implemented by processor 410) or hardware, or a combination of hardware and software. As a non-limiting example, the example can be implemented by one or more integrated circuits. Memory 420 can be any type suitable for the technical environment and can be implemented using any suitable data storage technology, such as optical memory devices, magnetic memory devices, semiconductor-based memory devices, fixed memory, and removable memory, as non-limiting examples. Processor 410 can be any type suitable for the technical environment and can include one or more microprocessors, general-purpose computers, special-purpose computers, and processors based on multi-core architectures, as non-limiting examples.

[0110] Various implementations involve decoding. As used herein, "decoding" can encompass, for example, all or part of the process performed on a received encoded sequence to produce a final output suitable for display. In various examples, such a process includes one or more processes typically performed by a decoder, such as entropy decoding, inverse quantization, inverse transform, and differential decoding. In various examples, such a process also or alternatively includes processes performed by a decoder in various implementations of this application, such as receiving a video image; receiving an indication of a ratio value corresponding to the video image; adjusting the number of codewords in a mapping function associated with the video image based on the ratio value; and decoding the video image based on the mapping function.

[0111] As further examples, in one example, "decoding" refers only to entropy decoding; in another example, "decoding" refers only to differential decoding; and in yet another example, "decoding" refers to a combination of entropy decoding and differential decoding. Whether the phrase "decoding process" is intended to specifically refer to a subset of operations or to refer to a broader decoding process will be determined according to the specific context of the description, and it is believed that those skilled in the art will understand well.

[0112] Various implementations involve encoding. Similar to the discussion of "decoding" above, the term "encoding" as used in this application can encompass all or part of the process performed, for example, on an input video sequence to produce an encoded bitstream. In various examples, such a process includes one or more procedures typically performed by an encoder, such as receiving a video image; receiving an indication of a ratio value corresponding to the video image; adjusting the number of codewords in a mapping function associated with the video image based on the ratio value; and encoding the video image based on the mapping function.

[0113] As further examples, in one example, "encoding" refers only to entropy coding; in another example, "encoding" refers only to differential coding; and in yet another example, "encoding" refers to a combination of differential and entropy coding. Whether the phrase "encoding process" is intended to specifically refer to a subset of operations or to refer to a broader encoding process will be determined according to the specific context of the description and is believed to be well understood by those skilled in the art.

[0114] It should be noted that the syntax elements used in this article (such as the coding syntax for intensity ranges, granular parameters, block offsets, scaling factors, etc.) are descriptive terms. Therefore, the use of other syntax element names is not excluded.

[0115] When a diagram is presented as a flowchart, it should be understood that it also provides a block diagram of the corresponding apparatus. Similarly, when a diagram is presented as a block diagram, it should be understood that it also provides a flowchart of the corresponding method / process.

[0116] The implementations and aspects described herein can be implemented, for example, in methods or processes, apparatuses, software programs, data streams, or signals. Even if discussed only in the context of a single form of implementation (e.g., discussed only as a method), implementations of the features discussed can be implemented in other forms (e.g., apparatuses or programs). Apparatuses can be implemented, for example, in suitable hardware, software, and firmware. Methods can be implemented, for example, in a processor, where processor refers generally to a processing device, including, for example, a computer, microprocessor, integrated circuit, or programmable logic device. Processors also include communication devices, such as computers, cellular phones, portable / personal digital assistants (PDAs), and other devices that facilitate information communication between end users.

[0117] References to "an example" or "example" or "one implementation" or "implementation" and their variations mean that a particular feature, structure, characteristic, etc., described in connection with that example is included in at least one example. Therefore, the phrases "in an example" or "in a sample" or "in an embodiment" or "in an implementation" appearing throughout this application, as well as any other variations, do not necessarily refer to the same example.

[0118] Furthermore, this application may involve "determining" various types of information. Determining information may include, for example, one or more of estimated information, calculated information, predicted information, or information retrieved from memory. Acquiring may include receiving, retrieving, constructing, generating, and / or determining.

[0119] Furthermore, this application may involve "accessing" various types of information. Accessing information may include, for example, receiving information, retrieving information (e.g., from memory), storing information, moving information, copying information, calculating information, determining information, predicting information, or estimating information, or one or more of these.

[0120] Furthermore, this application may relate to "receiving" various types of information. Like "accessing," "receiving" is intended to be a broad term. Receiving information may, for example, include access information or retrieval information (e.g., from memory) or one or more of them. Moreover, "receiving" generally refers in some way to operations such as storing information, processing information, sending information, moving information, copying information, erasing information, calculating information, determining information, predicting information, or estimating information.

[0121] It should be understood that the use of any of the following " / ", "and / or", and "...at least one", such as in "A / B", "A and / or B", and "at least one of A and B", is intended to cover selecting only the first listed option (A), or only the second listed option (B), or selecting both options (A and B). As a further example, in the cases of "A, B, and / or C" and "at least one of A, B, and C", such wording is intended to cover selecting only the first listed option (A), or only the second listed option (B), or only the third listed option (C), or only the first and second listed options (A and B), or only the first and third listed options (A and C), or only the second and third listed options (B and C), or selecting all three options (A, B, and C). It will be clear to those skilled in the art that this can be extended to the listed items.

[0122] Similarly, as used herein, the term "signal" refers, among other things, instructing the corresponding decoder to do something. Encoder signals may include, for example, the number of intensity intervals, the number of model values, particle parameters, particle identifiers, scaling factors, etc. Thus, in this example, the same parameters are used on both the encoder and decoder sides. Therefore, for example, the encoder can send a specific parameter to the decoder (explicit signaling) so that the decoder can use the same specific parameter. Conversely, if the decoder already has that specific parameter as well as other parameters, then signaling can be used without transmission (implicit signaling) to simply allow the decoder to know and select that specific parameter. Bit savings are achieved in various examples by avoiding the transmission of any actual function. It should be understood that signaling can be done in various ways. For example, in various examples, one or more syntax elements, flags, etc., are used to signal information to the corresponding decoder. While the foregoing refers to the verb form of the term "signal," the term "signal" can also be used as a noun in this document.

[0123] As will be apparent to those skilled in the art, implementations can generate signals in various formats to carry information, such as information that can be stored or transmitted. This information may include, for example, instructions for performing a method, or data generated by one of the described implementations. For example, the signal may be formatted to carry a bitstream of the described example. Such a signal may be formatted, for example, as an electromagnetic wave (e.g., using the radio frequency portion of the spectrum) or as a baseband signal. Formatting may include, for example, encoding a data stream and modulating a carrier wave with the encoded data stream. The information carried by the signal may be, for example, analog or digital information. As is known, the signal can be transmitted via a variety of different wired or wireless links. The signal may be stored on, or accessed or received from, a processor-readable medium.

[0124] This document describes several examples. Features of the examples may be provided individually or in any combination across various claim classes and types. Furthermore, examples may include one or more of the features, devices, or aspects described herein, individually or in any combination, across various claim classes and types. For example, features described herein may be implemented in a bitstream or signal that includes information generated as described herein. This information may allow a decoder to decode a bitstream, encoder, bitstream, and / or decoder according to any of the embodiments described. For example, features described herein may be implemented by creating and / or transmitting and / or receiving and / or decoding a bitstream or signal. For example, features described herein may be implemented as a method, process, apparatus, medium storing instructions, medium storing data, or signal. For example, features described herein may be implemented by a television, set-top box, cellular phone, tablet computer, or other electronic device performing decoding. The television, set-top box, cellular phone, tablet computer, or other electronic device may display (e.g., using a monitor, screen, or other type of display) a resulting image (e.g., an image reconstructed from the residual of a video bitstream). The television, set-top box, cellular phone, tablet computer, or other electronic device may receive a signal including an encoded image and perform decoding. The coding loop can be updated on both the encoder and decoder sides using a luminance mapping function based on reconstructed samples generated (e.g., already generated) in the current stripe or image. Compression efficiency may be affected. It is possible to reduce the bit rate while maintaining quality, and / or improve quality while maintaining the bit rate. This paper describes features associated with block partitioning.

[0125] The features described herein can be associated with Luminance Mapping and Chroma Scaling (LMCS) for SDR content. In some sequences (e.g., for captured content), the data range of luminance samples may or may not correspond to 1024 levels of a 10-bit signal. The data range of luminance samples may correspond to a range of 64–940 (e.g., a standard range). An LMCS function can modify the luminance of the content to fully utilize (e.g., 10-bit) capabilities. The mapping function can be applied to the luminance based on a pre-computed LUT, which is computed on the encoder side. The value of the LUT can be defined by the encoder based on the characteristics of the content. A 10-bit representation (e.g., 1024 levels) can be divided into 16 pieces (e.g., intervals), each with 64 levels. The derivation of the LMCS mapping function in the encoder may include the following steps in random access mode.

[0126] Video parameters can be calculated. For a video or video chunk, parameters can be extracted from the I-frame, such as the histogram and local variance of the frame brightness samples by segment. The result of the histogram can be given in 16 values ​​(corresponding to 16 segments). A segment can correspond to 1 / 16 of the value range of 0 to 1023 of a 10-bit signal. Hist[0] can correspond to the number of occurrences of samples with brightness between 0 and 63 on the (e.g., the entire) picture, Hist[1] corresponds to the number of occurrences of samples with brightness between 64 and 127, and so on. Hist[i] for i = 0 to 15 can be normalized such that Hist[0] + ... + Hist

[15] = 1.0.

[0127] The variance of a sample P with brightness values ​​P(x,y) can be calculated as 5. The variance is calculated by dividing the sum of the squares of the pixels within the window by the window size minus the square of the average value within the window. The variance can be summed with the variances of the same segments, and the segments can be determined by the brightness values ​​of the current sample.

[0128] It can be determined whether LMCS is applied. If both Hist[0] and Hist

[15] are below the threshold (e.g., 0.001), the LCMS function can be activated (e.g., codewords (CW) from 0-63 and 960-1023 can or can not be used to encode the sequence).

[0129] The global codeword count can be calculated. 896 codes are available to encode values ​​between 64 and 959. The global codeword count (NbCW) can be a multiple of 14. NbCW can be determined based on a threshold calculated from the histogram and variance values.

[0130] CW can be refined by segments. The number of CWs in each segment can be the same (e.g., NbCW can be a multiple of 14). The refinement process can use the histogram and variance values ​​of the segments to modify the number of CWs in each segment around the initial values ​​calculated (e.g., the segment for index i is denoted as m_binCW[i] below). The refinement process can include the following: for (int i = 0; i<16; i++){if (Hist[i]>0.001){ifHist[i]>0.4 { hist = 0.4};else {hist = Hist[i]};delta1 = (uint16_t)(10.0 hist + 0.5);delta2 = (uint16_t)(20.0 hist + 0.5);if (normVar[i]<0.8) {m_binCW[i] = m_binCW[i]+ delta2;}else if (normVar[i]<0.9){m_binCW[i] = m_binCW[i]+ delta1;}if (normVar[i]>1.2){m_binCW[i] = m_binCW[i]- delta2;}else if (normVar[i]>1.1) {m_binCW[i] = m_binCW[i]- delta1;}} Correction can be applied so that the sum of the CW numbers of each segment is not higher than 1023.

[0131] Examples can include QP <= 22. A process can be applied for QP <= 22 that uses a threshold to determine if NbCW has been modified. If NbCW has been modified, this value can be set to 924.

[0132] LUTs can be generated. In the example, LUTs corresponding to the number of cutaway frames (CWs) for each segment can be generated. The LUT corresponding to the number of CWs for each segment can be a Forward Mapping LUT and is used when transitioning from a non-mapped domain to a mapped domain. In parallel, Inverse Mapping LUTs can be generated based on the same number of CWs for each segment. Inverse Mapping LUTs can be used when transitioning from a mapped domain to a non-mapped domain. Forward Mapping and Inverse Mapping can be applied to frames of the video (e.g., frames other than I-frames).

[0133] LMCS parameters can be communicated using signals. LMCS parameters can be communicated by the encoder in the APS using signals.

[0134] Table 1: LMCS syntax table.

[0135]

[0136] The decoder can decode this information, reconstruct the forward mapping function and the inverse mapping function, and use these functions during the decoding process (for example, the forward mapping can be applied to predict the brightness samples between frames, while the inverse mapping can be applied after the reconstruction process and before the loop filtering).

[0137] The features described in this article can be associated with Luminance Mapping and Chroma Scaling (LMCS) for HDR content. In some SDR sequences (e.g., for captured content SDR sequences), the data range of the luminance signal may or may not correspond to 1024 levels (e.g., for 10-bit signals), and may correspond to a range of 64–940. For HDR signals, when the HDR signal is in PQ format, the code range may be reduced because a maximum value of 1023 may correspond to a peak luminance of 10,000 cd / m², which may or may not be used. Grading can use peak luminances of up to 1000, 4000, or 5000 cd / m², which may result in approximately 720, 850, and 875 luminance codewords, respectively. LMCS functions can modify the luminance signal of the content to fully utilize the 10-bit capability. Mapping functions can be applied to the luminance signal due to functions or LUTs pre-computed on the encoder side. The values ​​of the functions or LUTs can be defined by the encoder based on the characteristics of the content. A 10-bit representation (1024 levels) can be divided into a given number of N segments (e.g., 16 intervals), each segment containing 1024 / N levels (e.g., 64 when using 16 intervals). The luminance mapping function can be defined as a piecewise linear function, defined with N segments. For example, in the case of SDR, the Forward_Mapping function can be modeled by a piecewise linear model, for example, 16 segments.

[0138] For an input value x (e.g., in [0,1023] of a 10-bit signal), a Forward_Mapping LUT can be constructed from a piecewise linear function. In parallel, an Inverse_Mapping LUT can be generated based on the same piecewise linear function. An Inverse_Mapping LUT can be used when a transformation from the mapped domain to the unmapped domain is required. Forward_Mapping and Inverse_Mapping can be applied to images in video.

[0139] A mapping function can correspond to or be defined as a set of CW numbers for each segment, where a number of CW numbers (e.g., one) can be assigned to a segment of the function.

[0140] The encoder derives the LMCS function by extracting video parameters (e.g., histogram, variance) from I-frames and calculating the number of codewords (CWs) per segment based on the histogram and local variance. Since the number of CWs for each segment depends on the content, these values ​​can be calculated on the encoder side, and the number of CWs for each segment can be transmitted to the bitstream per I-frame. There may be (e.g., additional) costs associated with transmitting the data. The amount of data to be transmitted can be limited by selecting a finite set of predefined CWs for each segment (e.g., 4, 8, or 16). The encoder can determine the optimal set of CWs for each segment used for the current video segment (e.g., clip). As described herein, the optimal set of CWs for each segment used for the current segment can be referred to as the CODE.

[0141] A CODE can be selected by exhaustive rate-distortion optimization (RDO), which can involve encoding the image against possible CODEs and selecting the CODE that results in the optimal rate-distortion value. For a given CODE, full encoding of the image can be performed. Full encoding can select the optimal partitioning tree and encoding mode for the image's CTU among possible partitioning trees and available encoding modes, thus optimizing the rate-distortion criterion. The rate-distortion criterion can include a measure of the bit rate R(CODE) obtained when encoding the image (or video) and the distortion D(CODE) corresponding to the distance between samples of the decoded image and samples of the original image, as shown in Equation (1).

[0142] RD(CODE) = D(CODE) + λ R(CODE) (1) The parameter λ is a parameter (e.g., a Lagrange parameter) that can depend on the quantization parameter (QP) used to encode the image. D can correspond to the mean square error or the mean absolute error.

[0143] The encoding cost of signaling the CODE can be reduced, as described herein. An optimal set of "CWs for each segment of the current segment" (e.g., the CODE) can be selected, and this selection information can be indicated in the video data (e.g., transmitted to the bitstream). The process of estimating the CODE can be performed based on analysis of the input image to be encoded. The signaling mechanism can lead to a reduction in the signaling cost of the CODE.

[0144] CODEs can be identified based on input features computed from samples of the input image (or multiple images). These input features are used to boost the decision tree with gradients to select the best CODE from the set of possible CODEs.

[0145] Features calculated from input image samples can be based on the minimum and maximum values ​​of the video, histograms of image (or multiple image) samples, variance of image samples, time information, or parameters extracted from image (or multiple image) samples.

[0146] A possible set of codes can be obtained from an initial code by applying a fixed ratio to the value of the initial code. A possible set of codes can also be obtained from an initial code by changing at least one (e.g., one) value of the initial code. Features can be computed from the mapped input image samples using at least one (e.g., one) code from the possible set of codes. An exemplary process may include signaling information related to the code into the bitstream to reduce code signaling costs.

[0147] The features described in this paper can be associated with gradient boosting decision trees. The optimal code can be selected using a regression algorithm called LightGBM. LightGBM is a gradient boosting framework that uses a tree-based learning algorithm.

[0148] Once LightGBM training is complete, the goal of the inference process can be to predict the gain using a given code, given input features computed on the image. The set of input features used for both the training and inference processes can be defined.

[0149] The examples described in this article can be applied to various types of machine learning algorithms, such as random forests and neural networks that employ deep learning algorithms.

[0150] To train the LightGBM model, a dataset can be obtained by encoding the test sequence set with different codes at (e.g., different) QPs (e.g., QP 22, 27, 32, 37). The set of codes for selecting the best code can be defined by applying (e.g., different) a ratio to the number of codewords in the initial code, which may correspond to an identity mapping function or a predefined mapping function. The process of calculating the code with the ratio R may include: defining MIN and MAX values ​​for the current frame; defining the total number of segments N between the segments to which the MIN value belongs and the segments to which the MAX value belongs; and applying the ratio R to the number of codewords in the segments (e.g., the ratio R can satisfy condition N). 64 R<2 n ), where n can be the bit depth of the video sample.

[0151] In the example, the initial CODE can correspond to 64 CWs per segment for 16 segments of a 10-bit signal. If the minimum number of video samples is 64 and the maximum is 960, then N=14. Applying a ratio R of 1.1 means that the number of CWs per segment can be equal to ROUND(64). 1.1) = 70. In the example, the ratio R = 1.1 is acceptable because 14 70 = 980 < 1024.

[0152] In the exemplary process, the initial CODE can be represented as a LUT, and the elements of the LUT can be scaled as shown in equation (2): scaledLUT[x] = Min(2^n – 1, round(R) LUT[x]), where x = 0 to 2^n (2) where round(.) can return the nearest integer value.

[0153] For segment i, nbCWs(i) can be computed as (scaledLUT[(i+1)]) K]–scaledLUT[i K]), where K = 2^n / maxNbPieces, and maxNbPieces is the maximum number of segments (e.g., for a 10-bit signal, maxNbPieces = 16, K = 64).

[0154] For the sequence and the code, the coding gain and input feature values ​​can be derived. The coding gain can be the BD rate coding gain, the PSNR coding gain, or a probability of coding gain. The BD rate coding gain can provide an indication of the average bit rate change over 4 QPs between the anchor (e.g., the sequence encoded with the initial code) and the test (e.g., the sequence encoded with the code). Negative values ​​can indicate a decrease in bit rate, which can be associated with the coding gain. The PSNR coding gain can provide an indication of the average PSNR change over 4 QPs between the anchor (e.g., the sequence encoded with the initial code) and the test (e.g., the sequence encoded with the code). The probability of coding gain can provide the average probability of coding gain over 4 QPs between the anchor (e.g., the sequence encoded with the initial code) and the test (e.g., the sequence encoded with the code). The gain can be the BD rate coding gain, the PSNR coding gain, or a probability of coding gain. The cost function of the training process can be the L1 or L2 distance between the predicted gain and the actual gain.

[0155] Exemplary features can be used for training and inference. Input features for the LightGBM model can include the following: MIN and MAX values ​​of image samples; and / or histograms of image samples segmented. For a segment, the number of pixels belonging to that segment can be calculated as follows: For a piece NB_pixel[piece] = 0 For a pixel (between 0 and Nb_pixel_picture) For a piece (between 0 and 15) If piece 64 <pixel_value<(piece+1) `NB_pixel[piece] += 1` calculates the average variance of an image sample divided into segments. For a segment, the average variance of pixels belonging to that segment can be calculated as follows: For a piece, `Sum_variance[piece] = 0`, `NB_pixel[piece] = 0`. For a pixel (between 0 and Nb_pixel_picture), calculate the variance of the pixel. For a piece (between 0 and 15), if piece... 64 <pixel_value<(piece+1) 64Sum_variance[piece] +=variance of the pixelNB_pixel[piece] +=1For a piece (between 0 and 15)Mean_variance[piece] = Sum_variance[piece] / NB_pixel[piece] The temporal variance can be calculated between the current image Icur and the distant image Imc that has been motion compensated. In the example, the temporal variance between the current image Icur and the distant image with a temporal distance d can be calculated as follows, as shown in equation (3): (3) The image can be divided into blocks, N blk It refers to the number of blocks.

[0156] When applying the forward mapping function (fwdMap) and the backward mapping function (invMap) to image samples, the parameters (e.g., the sum of squared errors, and / or the number of non-zero errors) can be shown in equations (4) and (5): (4) (5) Due to rounding, the piecewise linear inverse mapping function may or may not be the exact inverse of the piecewise linear forward mapping function.

[0157] The characteristics of an image or video (e.g., any other characteristics) can be used as features for decision-making algorithms.

[0158] Figure 5 illustrates an exemplary flowchart for inference used in ratio selection. Figure 5 provides an exemplary flowchart of the inference process in an example of a CODE defined by a fixed ratio. The input to the process can be an image, such as the first image of the clip to be encoded or the first image of the segment to be encoded when the video is encoded in segments (e.g., when random access encoding encodes intra-frame images in a bitstream). At 501, the parameter maxGain can be set to 0. This parameter can store the maximum gain value among the possible ratio values ​​to be tested. At 502, a loop can be performed over the possible ratio values. At 503, the features under consideration can be computed. At 504, a machine learning model can be used to perform inference, with the features from 503 as input. Gain estimation can be performed when using the tested ratio values. The gain estimated at 504 can be compared with maxGain at 505. At 506, if the estimated gain is higher than maxGain, the maxGain value can be updated with that gain, and the current ratio can be stored as best_ratio. If the estimated gain is not higher than maxGain, the next step (e.g., 507) can be performed. At 507, it can be checked whether the end of the loop for possible ratio values ​​has been reached. If the end of the loop has not been reached, the next ratio value can be tested (e.g., at 502). If the end of the loop has been reached, the output of the inference process can be the estimated best ratio value, corresponding to the value best_ratio.

[0159] The number of codewords (CWs) for each segment can be defined individually for each segment. In the example, the set of CODEs for selecting the best CODE can be defined by modifying the number of codewords in at least one (e.g., one) segment of the initial CODE. The initial CODE can have an equal number of CWs in each segment. The process of calculating the CODEs belonging to the CODE set can include the following operations: defining MIN and MAX values ​​for the current frame; defining the total number of segments N between the segments to which the MIN value belongs and the segments to which the MAX value belongs; and / or modifying the number of codewords in at least one (e.g., one) segment of the initial CODE (e.g., the code set can satisfy a condition). (where n is the bit depth of the video sample). In the example, the initial CODE could correspond to 64 CWs per segment for 16 segments of a 10-bit signal. If the minimum value of the video is 64 and the maximum value is 960, then N=14, and 896 levels out of 1023 can be used. CODEs can be considered where at least one segment has more than 64 CWs. The set of CODEs can be defined in this way, and the optimal CODE for the sequence can be determined.

[0160] For the sequence and the CODE, the coding gain and the input feature value can be derived.

[0161] The cost function used for the training process can be defined as the L1 or L2 distance between the predicted gain and the actual gain. Figure 6 illustrates an exemplary flowchart for inference of a set of CODEs. Figure 6 shows an exemplary inference process in an example of selecting a CODE from a set of CODEs. The input to the process can be an image, such as the first image of the clip to be encoded or the first image of the segment to be encoded when the video is encoded in segments (e.g., when random access encoding encodes regular intra-frame images in a bitstream). At 601, the parameter maxGain can be set to 0. maxGain can store the maximum gain value among the possible CODEs to be tested. At 602, a loop can be performed on the possible CODEs. At 603, the features under consideration can be computed. At 604, inference can be performed using a machine learning model, with the features from 603 as input. At 603, gain estimation can be performed when using the CODE to be tested. The gain estimated from 604 can be compared with maxGain at 605. At 606, if the estimated gain is higher than maxGain, the maxGain value can be updated with that gain, and the current CODE can be stored as best_code. If the estimated gain is not higher than maxGain, the next step can be executed (e.g., at 607). At 607, it can be checked whether the end of the loop for possible ratio values ​​has been reached. If the end of the loop has not been reached, the next CODE can be tested (e.g., at 602). If the end of the loop has been reached, the output of the inference process can be the estimated best CODE, corresponding to the value best_code.

[0162] Features can be computed from the input image samples of the mapping using at least one (e.g., one) CODE from a possible set of CODEs. In the example, features can be extracted from a video (e.g., one to which a forward mapping function has been applied). In the example, when the features extracted from the content are based on a histogram, the number of samples belonging to a segment can be calculated as follows: For a piece NB_pixel[piece] = 0 For a pixel (between 0 and Nb_pixel_picture) For a piece (between 0 and 15) If piece 64 <Forward_mapping(pixel_value)<(piece+1) Figure 7 illustrates an exemplary flowchart of inference using features computed based on mapped video. Figure 7 illustrates an exemplary flowchart of the inference process when the CODE is defined by a fixed ratio. The input to the process can be a picture, such as the first picture of the clip to be encoded or the first picture of the segment to be encoded when the video is encoded in segments (e.g., when random access encoding encodes regular intra-frame pictures in a bitstream). At 701, the parameter maxGain can be set to 0. maxGain can store the maximum gain value among the possible ratio values ​​to be tested. At 702, a loop can be performed over the possible ratio values. At 703, the considered features can be computed on the mapped video corresponding to (e.g., all) ratios. At 704, inference can be performed using a machine learning model, with the features from 703 as input. Gain estimation can be performed when using the tested ratio values. The gain estimated from 704 can be compared with maxGain at 705. In step 706, if the estimated gain is higher than `maxGain`, the `maxGain` value can be updated using that gain, and the best ratio can be stored as `best_ratio`. If the estimated gain is not higher than `maxGain`, the next step can be executed (e.g., in step 707). In step 707, it can be checked whether the end of the loop for possible ratio values ​​has been reached. If the end of the loop has not been reached, the next ratio value can be tested (e.g., in step 702). If the end of the loop has been reached, the output of the inference process can be the estimated best ratio value, corresponding to the value `best_ratio`.

[0163] Information related to the optimal CODE can be indicated in the video data (e.g., signaled to the bitstream). When LMCS is applied, the lmcs_data set can be signaled (e.g., on the encoder side) or parsed (e.g., on the decoder side), and indications (e.g., flags) can be added to the LMCS data. For example, the indicator lmcs_cw_code_selection_flag can indicate whether an LMCS codeword is transmitted individually or belongs to a predefined set of codes. When lmcs_cw_code_selection_flag is false, a codeword can be transmitted. When lmcs_cw_code_selection_flag is true, lmcs_cw_ratio_code_flag can be tested. When the indicator lmcs_cw_ratio_code_flag is true, the ratio lmcs_cw_ratio can be transmitted. When lmcs_cw_ratio_code_flag is false, the index lmcs_cw_nb_codes of the predefined CODEs (e.g., within a predefined set of CODEs) can be transmitted.

[0164] Table 2: LMCS Syntax Table

[0165] In an example, the procedure could include the following: If LMCS is enabled for the picture (ph_lmcs_enabled_flag is true)Execute lmcs_dataIf lmcs_cw_code_selection_flag is falseRecompute LMCS codewords fromlmcs_delta_abs_cw[i] and lmcs_delta_sign_cw_flag[i]Else if lmcs_cw_code_selection_flag is trueIf lmcs_cw_ratio_code_flag is trueRecompute LMCS codewords from the lmcs_cw_ratio and the initial valueof codewords (64).Else If lmcs_cw_ratio_code_flag is falseGet associated LMCS codewords to the corresponding code number lmcs_cw_code_idx (up to 16 codes with 4 bits, or 4 codes if 2 In the example, one of lmcs_cw_ratio and lmcs_cw_code_idx can be signaled. In the example, lmcs_cw_ratio can correspond to the index used to retrieve the ratio value from a predefined ratio table.

[0166] While the features and elements have been described above in specific combinations, those skilled in the art will understand that each feature or element can be used alone or in combination with any other features and elements. Furthermore, the methods described herein can be implemented in computer programs, software, or firmware, which are incorporated in a computer-readable medium for execution by a computer or processor. Examples of computer-readable media include electronic signals (transmitted via wired or wireless connections) and computer-readable storage media. Examples of computer-readable storage media include, but are not limited to, read-only memory (ROM), random access memory (RAM), registers, cache memory, semiconductor memory devices, magnetic media such as internal hard disks and removable disks, magneto-optical media, and optical media such as CD-ROMs and digital versatile optical discs (DVDs). The processor associated with the software can be used to implement a radio frequency transceiver used in a WTRU, UE, terminal, base station, RNC, or any host computer.

Claims

1. A video encoding device, comprising: A processor configured to: determine a plurality of candidate ratio values ​​associated with video images; and determine a plurality of corresponding coding gains associated with the plurality of candidate ratio values; A ratio value is selected from the plurality of candidate ratio values ​​based on the coding gain associated with each of the candidate ratio values; the number of codewords in the mapping function associated with the video image is adjusted based on the selected ratio value; and the video image is encoded based on the mapping function.

2. The video encoding apparatus of claim 1, wherein the processor is further configured to: divide the range of luminance sample values ​​of the video image into a plurality of luminance sample value segments, wherein the plurality of luminance sample value segments correspond to corresponding plurality of luminance value sub-ranges; determine the number of codewords for the plurality of luminance sample value segments based on a selected ratio value; map the plurality of luminance sample value segments to updated luminance sample values ​​based on the number of codewords determined for the plurality of luminance sample value segments; and generate the mapping function based on the mapped plurality of luminance sample value segments.

3. The video encoding apparatus of any one of claims 1 and 2, wherein the processor is further configured to include an indication of a selected ratio value in the video data, wherein the video picture is further encoded based on the indication.

4. The video encoding apparatus of claim 3, wherein the processor is further configured to: determine whether to signal a plurality of codewords individually or to associate the plurality of codewords with a predefined code, wherein, based on the determination that the plurality of codewords are associated with the predefined code, the indication of a selected ratio value is included in the video data.

5. The video encoding apparatus of any one of claims 1 and 2, wherein the processor is further configured to: determine that a codeword is associated with a predefined code; and include an index associated with the codeword in the video data, wherein the video images are further encoded based on the codeword.

6. The video encoding apparatus of any one of claims 1 to 5, wherein the processor is further configured to: determine at least one feature from a plurality of mapped video picture samples associated with the video picture using at least one ratio value from the plurality of candidate ratio values, wherein the encoding gain is determined based on the determined at least one feature.

7. The video encoding apparatus of claim 6, wherein the at least one feature includes at least one of the following: the minimum value of the plurality of mapped video image samples, the maximum value of the plurality of mapped video image samples, a segmented histogram of the plurality of mapped video image samples, a segmented variance of the plurality of mapped video image samples, and the temporal variance between the plurality of mapped video image samples and a sample of a second video image different from the video image.

8. The video encoding apparatus of any one of claims 1 to 7, wherein the processor is further configured to: apply a gradient boosting decision tree algorithm to the plurality of candidate ratio values, wherein the encoding gain is determined based on applying the gradient boosting decision tree algorithm to the plurality of candidate ratio values.

9. A video decoding device, comprising: A processor configured to: receive a video image; receive an indication of a ratio value corresponding to the video image; adjust the number of codewords in a mapping function associated with the video image based on the ratio value; and decode the video image based on the mapping function.

10. The video decoding device of claim 9, wherein the processor is further configured to: divide the range of brightness sample values ​​of the video image into multiple brightness sample value segments, wherein the multiple brightness sample value segments correspond to corresponding multiple brightness value sub-ranges; Based on the ratio value, the number of codewords is determined for segments of the plurality of luminance sample values; based on the number of codewords determined for segments of the plurality of luminance sample values, the plurality of luminance sample values ​​are mapped to updated luminance sample values. The mapping function is generated based on segments of multiple mapped luminance sample values.

11. The video decoding apparatus of any one of claims 9 to 10, wherein the processor is further configured to: apply the mapping function to a plurality of inter-frame predicted luminance samples; and further decode the video image based on applying the mapping function to the plurality of inter-frame predicted luminance samples.

12. The video decoding apparatus of any one of claims 8 to 11, wherein the processor is further configured to: determine that a plurality of codewords belong to predefined codewords, wherein the indication of the ratio value is received based on the determination that the plurality of codewords belong to predefined codewords.

13. The device of any one of claims 1 to 12, further comprising a memory operatively connected to the processor.

14. A method for video encoding, the method comprising: Identify multiple candidate ratio values ​​associated with video images; Determine a plurality of corresponding coding gains associated with the plurality of candidate ratio values; A ratio value is selected from the plurality of candidate ratio values ​​based on the coding gain associated with each of the candidate ratio values; the number of codewords in the mapping function associated with the video image is adjusted based on the selected ratio value; and the video image is encoded based on the mapping function.

15. The method of claim 14, wherein the method further comprises: The range of brightness sample values ​​of the video image is divided into multiple brightness sample value segments, wherein the multiple brightness sample value segments correspond to corresponding multiple brightness value sub-ranges; Based on the selected ratio value, the number of codewords is determined for segments of the plurality of luminance sample values; based on the number of codewords determined for segments of the plurality of luminance sample values, the plurality of luminance sample values ​​are mapped to updated luminance sample values. The mapping function is generated based on segments of multiple mapped luminance sample values.

16. The method of any one of claims 14 and 15, wherein the method further comprises: The video data includes an indication of the selected ratio value, wherein the video images are further encoded based on the indication.

17. The method of claim 16, wherein the method further comprises: Determine whether to notify multiple codewords individually or whether the multiple codewords are associated with a predefined code, wherein, based on the determination that the multiple codewords are associated with a predefined code, the indication of the selected ratio value is included in the video data.

18. The method of any one of claims 14 and 15, wherein the method further comprises: The codeword is associated with a predefined code; And the video data includes an index associated with the codeword, wherein the video images are further encoded based on the codeword.

19. The method of any one of claims 14 to 18, wherein the method further comprises: At least one feature is determined from a plurality of mapped video image samples associated with the video image using at least one ratio value from the plurality of candidate ratio values, wherein the coding gain is determined based on the determined at least one feature.

20. The method of claim 19, wherein the at least one feature includes at least one of the following: the minimum value of the plurality of mapped video image samples, the maximum value of the plurality of mapped video image samples, a segmented histogram of the plurality of mapped video image samples, a segmented variance of the plurality of mapped video image samples, and the temporal variance between the plurality of mapped video image samples and a sample of a second video image different from the video image.

21. The method of any one of claims 14 to 20, wherein the method further comprises: The gradient boosting decision tree algorithm is applied to the plurality of candidate ratio values, wherein the encoding gain is determined based on applying the gradient boosting decision tree algorithm to the plurality of candidate ratio values.

22. A method for video decoding, the method comprising: Receive videos and images; Receive an indication corresponding to the ratio value of the video image; Based on the ratio value, the number of codewords in the mapping function associated with the video image is adjusted; and the video image is decoded based on the mapping function.

23. The method of claim 22, wherein the method further comprises: The range of brightness sample values ​​of the video image is divided into multiple brightness sample value segments, wherein the multiple brightness sample value segments correspond to corresponding multiple brightness value sub-ranges; Based on the ratio value, the number of codewords is determined for segments of the plurality of luminance sample values; based on the number of codewords determined for segments of the plurality of luminance sample values, the plurality of luminance sample values ​​are mapped to updated luminance sample values. The mapping function is generated based on segments of multiple mapped luminance sample values.

24. The method of any one of claims 22 to 23, wherein the method further comprises: The mapping function is applied to multiple brightness samples predicted by inter-frame analysis; Furthermore, the video images are decoded by applying the mapping function to the plurality of inter-frame predicted luminance samples.

25. The method of any one of claims 22 to 24, wherein the method further comprises: Determine that multiple codewords belong to predefined codewords, wherein the indication of the ratio value is received based on the determination that the multiple codewords belong to predefined codewords.

26. A computer-readable medium comprising instructions for causing one or more processors to perform the method as described in any one of claims 14 to 25.

27. A computer program product stored on a non-transitory computer-readable medium and comprising program code instructions for implementing the steps of the method according to at least one of claims 14 to 25 when executed by a processor.

28. Video data comprising information representing a video image encoded according to any one of claims 14 to 21.