Decoder-side intra mode derivation merging

By using distance and scaling factors to scale the gradient histogram in the intra-frame mode derivation merging mode on the decoder side, and mixing based on fusion weights, the inefficiency problem in the existing technology is solved, and more efficient video encoding and decoding is achieved.

CN121925840APending Publication Date: 2026-04-24INTERDIGITAL CE PATENT HOLDINGS SAS
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
INTERDIGITAL CE PATENT HOLDINGS SAS
Filing Date
2024-09-27
Publication Date
2026-04-24

Smart Images

  • Figure CN121925840A_ABST
    Figure CN121925840A_ABST
Patent Text Reader

Abstract

Systems, methods, and apparatus are disclosed for encoding and decoding in a decoder-side intra mode derivation (DIMD) merge mode. An example device may determine a first reference sample associated with a first gradient histogram and a second reference sample associated with a second gradient histogram. The apparatus may scale the first gradient histogram based on a distance between the first reference sample and a current block, and scale the second gradient histogram based on a distance between the second reference sample and the current block. The apparatus may derive a combined gradient histogram based on the scaled first and second gradient histograms, and derive an intra prediction mode based on the combined gradient histogram. The apparatus may obtain a prediction of the current block based on the intra prediction mode, and decode the current block based on the prediction.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-reference to related applications This application claims the benefit of European Provisional Application No. EP23306627.3, filed on 28 September 2023, the contents of which are incorporated herein by reference. Background Technology

[0002] Video coding systems can be used to compress digital video signals, for example, to reduce the storage and / or transmission bandwidth required for such signals. Video coding systems can include, for example, block-based, wavelet-based, and / or object-based systems. Summary of the Invention

[0003] Systems, methods, and instruments for encoding and / or decoding in intra-frame mode derivation merging modes on the decoder side are disclosed. Example devices may have a processor configured to perform one or more actions.

[0004] An example device (e.g., for video decoding) can determine whether to enable a decoder-side intra-frame mode derivation (DIMD) merging mode for a current block. The device can determine reference samples to be used for the DIMD merging mode. The reference samples may include a first reference sample associated with a first gradient histogram and a second reference sample associated with a second gradient histogram. The device can scale the first gradient histogram based on a first distance between the first reference sample and the current block. The device can scale the second gradient histogram based on a second distance between the second reference sample and the current block. The device can derive a merged gradient histogram based on the scaled first gradient histogram and the scaled second gradient histogram. The device can decode the current block based on the merged gradient histogram.

[0005] The device can determine a first scaling factor based on a first distance between the first reference sample and the current block. The device can then scale the first gradient histogram using the first scaling factor. The device can also determine a second scaling factor based on a second distance between the second reference sample and the current block. The device can then scale the second gradient histogram using the second scaling factor.

[0006] The first scaling factor can be inversely proportional to the first distance. The second scaling factor can be inversely proportional to the second distance.

[0007] The reference sample may include a third reference sample associated with the first gradient histogram. The device may determine a third distance between the third reference sample and the current block. The device may add or average the first distance and the third distance to determine a fourth distance. The device may expand or shrink the first gradient histogram based on the fourth distance.

[0008] The device can determine a first fusion weight based on a first bar amplitude associated with a scaled first gradient histogram, and a second fusion weight based on a second bar amplitude associated with a scaled second gradient histogram. The device can blend the scaled first gradient histogram and the scaled second gradient histogram based on the first fusion weight and the second fusion weight.

[0009] The device can derive one or more intra-frame prediction modes based on a merged gradient histogram. The device can obtain a prediction for the current block based on the one or more intra-frame prediction modes. The device can decode the current block based on the prediction.

[0010] An example device (e.g., for video encoding) can determine whether to enable a decoder-side intra-frame mode derivation (DIMD) merging mode for a current block. The device can determine reference samples to be used for the DIMD merging mode. The reference samples may include a first reference sample associated with a first gradient histogram and a second reference sample associated with a second gradient histogram. The device can scale the first gradient histogram based on a first distance between the first reference sample and the current block. The device can scale the second gradient histogram based on a second distance between the second reference sample and the current block. The device can derive a merged gradient histogram based on the scaled first gradient histogram and the scaled second gradient histogram. The device can encode the current block based on the merged gradient histogram.

[0011] The device can determine a first scaling factor based on a first distance between the first reference sample and the current block. The device can then scale the first gradient histogram using the first scaling factor. The device can also determine a second scaling factor based on a second distance between the second reference sample and the current block. The device can then scale the second gradient histogram using the second scaling factor.

[0012] The first scaling factor can be inversely proportional to the first distance. The second scaling factor can be inversely proportional to the second distance.

[0013] The reference sample may include a third reference sample associated with the first gradient histogram. The device may determine a third distance between the third reference sample and the current block. The device may add or average the first distance and the third distance to determine a fourth distance. The device may expand or shrink the first gradient histogram based on the fourth distance.

[0014] The device can determine a first fusion weight based on a first bar amplitude associated with a scaled first gradient histogram, and a second fusion weight based on a second bar amplitude associated with a scaled second gradient histogram. The device can blend the scaled first gradient histogram and the scaled second gradient histogram based on the first fusion weight and the second fusion weight.

[0015] The device can derive one or more intra-frame prediction modes based on a merged gradient histogram. The device can obtain a prediction for the current block based on the one or more intra-frame prediction modes. The device can decode the current block based on the prediction.

[0016] The device (e.g., a decoder, such as a video decoder) can determine whether to enable a decoder-side intra-frame mode derivation (DIMD) merging mode for the current block. The device can determine a reference region to be used for the DIMD merging mode. The device can decode the current block based on the reference region.

[0017] The reference region may include one or more of the following: the left region, top region, or upper-left region adjacent to the current block. The reference region may be determined based on the DIMD merge mode indication.

[0018] Decoding the current block based on the reference region may involve: identifying multiple gradient histograms associated with a reconstructed block encoded in DIMD or DIMD merging mode in the reference region; deriving a merged gradient histogram based on the multiple gradient histograms; deriving one or more intra-prediction modes based on the merged gradient histogram; obtaining a prediction of the current block based on the one or more intra-prediction modes; and decoding the current block based on the prediction.

[0019] An example device (e.g., for video decoding) can determine whether to enable a decoder-side intra-frame mode derivation (DIMD) merging mode for a current block. The device can determine a reference region to be used for the DIMD merging mode. The reference region may include multiple reference samples. Each reference sample may be associated with a corresponding gradient histogram. The device can scale each of the corresponding gradient histograms based on the distance between the associated reference sample and the current block to generate a scaled gradient histogram. The device can derive a merged gradient histogram based on the scaled gradient histogram. The device can decode the current block based on the merged gradient histogram.

[0020] An example device (e.g., for video decoding) can identify multiple gradient histograms associated with a reconstructed block encoded in a DIMD or DIMD merging mode. The device can derive multiple intra-prediction modes based on the multiple gradient histograms. The device can derive DIMD fusion weights for each of the intra-prediction modes based on the corresponding bar amplitude associated with each intra-prediction mode. The device can construct a prediction based on (one or more) intra-prediction modes and the DIMD fusion weights. The device can decode the current block based on the prediction.

[0021] The device can store one or more intra-prediction patterns and can decode a second block based on the stored intra-prediction patterns. Attached Figure Description

[0022] Figure 1A This is a system diagram illustrating an example communication system in which one or more of the disclosed embodiments may be implemented.

[0023] Figure 1B The illustration shows a device that can be used according to one embodiment. Figure 1A The diagram shows a system diagram of an example wireless transmit / receive unit (WTRU) used in a communication system.

[0024] Figure 1C The illustration shows a device that can be used according to one embodiment. Figure 1A The diagram shows a system diagram of an example radio access network (RAN) and an example core network (CN) used in a communication system.

[0025] Figure 1D The illustration shows a device that can be used according to one embodiment. Figure 1A The diagram shows a system diagram of a further example RAN and a further example CN used within the communication system.

[0026] Figure 2 The illustration shows a sample video encoder.

[0027] Figure 3 The illustration shows an example video decoder.

[0028] Figure 4 The illustration shows an example of a system that can implement various aspects and examples.

[0029] Figure 5 The illustration shows an example L-shaped template around the current block (e.g., coding unit (CU)).

[0030] Figure 6 An example technique for deriving prediction blocks is illustrated.

[0031] Figure 7 The diagram illustrates the example decoder-side intra-frame mode derivation (DIMD) process.

[0032] Figure 8 The example reference areas for DIMD, DIMD_T, and DIMD_L are illustrated.

[0033] Figure 9 The illustration shows an example DIMD process based on the merged gradient histogram (MHoG).

[0034] Figure 10 The illustration shows an example of adaptive reference region DIMD merging.

[0035] Figure 11 The illustration shows example adjacent blocks encoded in DIMD or DIMD merge in the left, top-left, and top regions.

[0036] Figure 12 The illustration shows example adjacent blocks encoded in DIMD or DIMD merge in the left, top-left, and top regions and their reference regions. Detailed Implementation

[0037] A more detailed understanding can be obtained from the following description, which is given in conjunction with the accompanying drawings as examples.

[0038] Figure 1AThis diagram illustrates an example communication system 100 in which one or more of the disclosed embodiments may be implemented. The communication system 100 may be a multiple access system that provides content (such as voice, data, video, messaging, broadcasting, etc.) to multiple wireless users. The communication system 100 enables multiple wireless users to access such content by sharing system resources (including wireless bandwidth). For example, the communication system 100 may employ one or more channel access methods, such as Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Frequency Division Multiple Access (FDMA), Orthogonal FDMA (OFDMA), Single Carrier FDMA (SC-FDMA), Zero-Tail Unique Word DFT Extended OFDM (ZT UW DTS-s OFDM), Unique Word OFDM (UW-OFDM), Resource Block Filtered OFDM, Filter Bank Multicarrier (FBMC), etc.

[0039] like Figure 1A As shown, the communication system 100 may include wireless transmit / receive units (WTRUs) 102a, 102b, 102c, 102d, RAN 104 / 113, CN 106 / 115, Public Switched Telephone Network (PSTN) 108, Internet 110, and other networks 112. Although it will be appreciated, the disclosed embodiments contemplate any number of WTRUs, base stations, networks, and / or network elements. Each of the WTRUs 102a, 102b, 102c, and 102d can be any type of device configured to operate and / or communicate in a wireless environment. As an example, WTRUs 102a, 102b, 102c, and 102d (any of which may be referred to as a “station” and / or “STA”) may be configured to transmit and / or receive wireless signals and may include user equipment (UE), mobile stations, fixed or mobile subscriber units, subscription-based units, pagers, cellular phones, personal digital assistants (PDAs), smartphones, laptops, netbooks, personal computers, wireless sensors, hotspots or Mi-Fi devices, Internet of Things (IoT) devices, watches or other wearable devices, head-mounted displays (HMDs), vehicles, drones, medical devices and applications (e.g., remote surgery), industrial devices and applications (e.g., robots and / or other wireless devices operating in industrial and / or automated processing chain scenarios), consumer electronics devices, devices operating on commercial and / or industrial wireless networks, etc. Any of WTRUs 102a, 102b, 102c, and 102d may be interchangeably referred to as a UE.

[0040] The communication system 100 may also include base station 114a and / or base station 114b. Each of base stations 114a and 114b may be any type of device configured to wirelessly interface with at least one of WTRUs 102a, 102b, 102c, and 102d to facilitate access to one or more communication networks, such as CN 106 / 115, the Internet 110, and / or other networks 112. As an example, base stations 114a and 114b may be any of a base transceiver station (BTS), Node-B, eNode B, home node B, home eNode B, gNB, NR NodeB, site controller, access point (AP), wireless router, etc. Although base stations 114a and 114b are depicted as single elements, it will be understood that base stations 114a and 114b may include any number of interconnected base stations and / or network elements.

[0041] Base station 114a may be part of RAN 104 / 113, which may also include other base stations and / or network elements (not shown), such as base station controllers (BSCs), radio network controllers (RNCs), relay nodes, etc. Base station 114a and / or base station 114b may be configured to transmit and / or receive radio signals on one or more carrier frequencies, which may be referred to as cells (not shown). These frequencies may be licensed spectrum, unlicensed spectrum, or a combination of licensed and unlicensed spectrum. A cell may provide coverage for a specific geographic area for a radio service, which may be relatively fixed or may change over time. A cell may be further divided into cell sectors. For example, the cell associated with base station 114a may be divided into three sectors. Therefore, in one embodiment, base station 114a may include three transceivers, i.e., one transceiver for each sector of the cell. In one embodiment, base station 114a may employ multiple-input multiple-output (MIMO) technology, and multiple transceivers may be used for each sector of the cell. For example, beamforming can be used to transmit and / or receive signals in a desired spatial direction.

[0042] Base stations 114a and 114b can communicate with one or more of WTRUs 102a, 102b, 102c, and 102d via air interface 116, which can be any suitable wireless communication link (e.g., radio frequency (RF), microwave, centimeter wave, millimeter wave, infrared (IR), ultraviolet (UV), visible light, etc.). Air interface 116 can be established using any suitable radio access technology (RAT).

[0043] More specifically, as noted above, communication system 100 can be a multiple access system and can employ one or more channel access schemes, such as CDMA, TDMA, FDMA, OFDMA, SC-FDMA, etc. For example, base station 114a in RAN 104 / 113, and WTRUs 102a, 102b, and 102c can implement radio technologies such as Universal Mobile Telecommunications System (UMTS) Terrestrial Radio Access (UTRA), which can use Wideband CDMA (WCDMA) to establish air interfaces 115 / 116 / 117. WCDMA can include communication protocols such as High-Speed ​​Packet Access (HSPA) and / or Evolved HSPA (HSPA+). HSPA can include High-Speed ​​Downlink (DL) Packet Access (HSDPA) and / or High-Speed ​​UL Packet Access (HSUPA).

[0044] In one embodiment, base station 114a and WTRUs 102a, 102b, 102c can implement radio technologies such as Evolved UMTS Terrestrial Radio Access (E-UTRA), which can use Long Term Evolution (LTE) and / or Advanced LTE (LTE-A) and / or Advanced LTE Pro (LTE-A Pro) to establish air interface 116.

[0045] In one embodiment, base station 114a and WTRUs 102a, 102b, 102c can implement radio technologies such as NR radio access, which can use a new radio (NR) to establish an air interface 116.

[0046] In one embodiment, base station 114a and WTRUs 102a, 102b, and 102c can implement multiple radio access technologies. For example, base station 114a and WTRUs 102a, 102b, and 102c can jointly implement LTE radio access and NR radio access, for example, using the dual connectivity (DC) principle. Therefore, the air interface utilized by WTRUs 102a, 102b, and 102c can be characterized by multiple types of radio access technologies and / or transmissions sent to / from multiple types of base stations (e.g., eNBs and gNBs).

[0047] In other embodiments, base station 114a and WTRUs 102a, 102b, 102c can implement the following radio technologies, such as IEEE 802.11 (i.e., WiFi), IEEE 802.16 (i.e., WiMAX), CDMA2000, CDMA2000 1X, CDMA2000 EV-DO, Provisional Standard 2000 (IS-2000), Provisional Standard 95 (IS-95), Provisional Standard 856 (IS-856), Global System for Mobile Communications (GSM), Enhanced Data Rate GSM Evolution (EDGE), GSMEDGE (GERAN), etc.

[0048] Figure 1A Base station 114b can be, for example, a wireless router, a home node B, a home eNode B, or an access point, and can utilize any suitable RAT to facilitate wireless connectivity in a local area, such as a commercial area, home, vehicle, campus, industrial facility, air corridor (e.g., for drone use), road, etc. In one embodiment, base station 114b and WTRUs 102c, 102d can implement radio technologies such as IEEE 802.11 to establish a wireless local area network (WLAN). In one embodiment, base station 114b and WTRUs 102c, 102d can implement radio technologies such as IEEE 802.15 to establish a wireless personal area network (WPAN). In yet another embodiment, base station 114b and WTRUs 102c, 102d can utilize a cellular-based RAT (e.g., WCDMA, CDMA2000, GSM, LTE, LTE-A, LTE-A Pro, NR, etc.) to establish a picocell or femtocell. Figure 1A As shown, base station 114b may have a direct connection to Internet 110. Therefore, base station 114b may not be required to access Internet 110 via CN 106 / 115.

[0049] RAN 104 / 113 can communicate with CN 106 / 115, which can be any type of network configured to provide voice, data, application, and / or Voice over Internet Protocol (VoIP) services to one or more of WTRUs 102a, 102b, 102c, and 102d. Data may have different Quality of Service (QoS) requirements, such as different throughput requirements, latency requirements, error tolerance requirements, reliability requirements, data throughput requirements, mobility requirements, etc. CN 106 / 115 can provide call control, billing services, location-based services, prepaid calling, internet connectivity, video distribution, and / or perform advanced security functions, such as user authentication. Although... Figure 1AAlthough not shown, it will be understood that RAN104 / 113 and / or CN106 / 115 can communicate directly or indirectly with other RANs that use the same RAT as or a different RAT than RAN 104 / 113. For example, in addition to being connected to RAN 104 / 113, which can utilize NR radio technology, CN106 / 115 can also communicate with another RAN (not shown) that uses GSM, UMTS, CDMA 2000, WiMAX, E-UTRA, or WiFi radio technology.

[0050] CN 106 / 115 may also act as a gateway for WTRU 102a, 102b, 102c, 102d to access PSTN 108, the Internet 110, and / or other networks 112. PSTN 108 may include a circuit-switched telephone network providing Common Old-Style Telephone Service (POTS). The Internet 110 may include a global system of interconnected computer networks and devices using common communication protocols such as Transmission Control Protocol (TCP), User Datagram Protocol (UDP), and / or Internet Protocol (IP) from the TCP / IP Internet Protocol suite. Network 112 may include wired and / or wireless communication networks owned and / or operated by other service providers. For example, network 112 may include another CN connected to one or more RANs, which may use the same RAT as RAN 104 / 113 or a different RAT.

[0051] Some or all of the WTRUs 102a, 102b, 102c, and 102d in communication system 100 may include multi-mode capabilities (e.g., WTRUs 102a, 102b, 102c, and 102d may include multiple transceivers for communicating with different wireless networks via different wireless links). For example, Figure 1A The WTRU 102c shown can be configured to communicate with a base station 114a that can use cellular-based radio technology and a base station 114b that can use IEEE 802 radio technology.

[0052] Figure 1B This is a system diagram illustrating example WTRU 102. (Example:) Figure 1B As shown, WTRU 102 may include a processor 118, a transceiver 120, a transmit / receive element 122, a speaker / microphone 124, a keypad 126, a display / touchpad 128, non-removable memory 130, removable memory 132, a power supply 134, a global positioning system (GPS) chipset 136, and / or other peripherals 138, etc. It will be appreciated that WTRU 102 may include any sub-combination of the above-described elements while remaining consistent with the embodiments.

[0053] Processor 118 may be a general-purpose processor, a special-purpose processor, a conventional processor, a digital signal processor (DSP), multiple microprocessors, one or more microprocessors associated with a DSP core, a controller, a microcontroller, an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) circuit, any other type of integrated circuit (IC), a state machine, etc. Processor 118 may perform signal encoding, data processing, power control, input / output processing, and / or any other functions that enable WTRU 102 to operate in a wireless environment. Processor 118 may be coupled to transceiver 120, which may be coupled to transmitting / receiving element 122. Although... Figure 1B The processor 118 and transceiver 120 are depicted as separate components, but it will be understood that the processor 118 and transceiver 120 can be integrated together in an electronic package or chip.

[0054] Transmitting / receiving element 122 can be configured to transmit signals to or receive signals from a base station (e.g., base station 114a) via air interface 116. For example, in one embodiment, transmitting / receiving element 122 can be an antenna configured to transmit and / or receive RF signals. In one embodiment, transmitting / receiving element 122 can be a transmitter / detector configured to transmit and / or receive, for example, IR, UV, or visible light signals. In yet another embodiment, transmitting / receiving element 122 can be configured to transmit and / or receive both RF and optical signals. It will be appreciated that transmitting / receiving element 122 can be configured to transmit and / or receive any combination of wireless signals.

[0055] Although the transmitting / receiving element 122 is in Figure 1B While depicted as a single element, WTRU 102 may include any number of transmitting / receiving elements 122. More specifically, WTRU 102 may employ MIMO technology. Thus, in one embodiment, WTRU 102 may include two or more transmitting / receiving elements 122 (e.g., multiple antennas) for transmitting and receiving wireless signals via air interface 116.

[0056] Transceiver 120 can be configured to modulate signals to be transmitted by transmitting / receiving element 122 and demodulate signals received by transmitting / receiving element 122. As noted above, WTRU 102 can have multi-mode capability. Thus, for example, transceiver 120 may include multiple transceivers for enabling WTRU 102 to communicate via multiple RATs (such as NR and IEEE 802.11).

[0057] The processor 118 of WTRU 102 can be coupled to the speaker / microphone 124, keypad 126, and / or display / touchpad 128 (e.g., a liquid crystal display (LCD) unit or an organic light-emitting diode (OLED) display unit), and can receive user input data from them. The processor 118 can also output user data to the speaker / microphone 124, keypad 126, and / or display / touchpad 128. Additionally, the processor 118 can access information from any type of suitable memory (such as non-removable memory 130 and / or removable memory 132), and store data in that memory. Non-removable memory 130 may include random access memory (RAM), read-only memory (ROM), hard disk, or any other type of memory storage device. Removable memory 132 may include a subscriber identity module (SIM) card, memory stick, secure digital storage (SD) card, etc. In other embodiments, processor 118 may access information from memory that is not physically located on WTRU 102 (such as on a server or home computer (not shown)) and store data in that memory.

[0058] The processor 118 can receive power from the power supply 134 and can be configured to distribute and / or control the power going to other components in the WTRU 102. The power supply 134 can be any suitable device for powering the WTRU 102. For example, the power supply 134 may include one or more dry cell batteries (e.g., nickel-cadmium (NiCd), nickel-zinc (NiZn), nickel metal hydride (NiMH), lithium-ion (Li-ion), etc.), solar cells, fuel cells, etc.

[0059] The processor 118 may also be coupled to a GPS chipset 136, which may be configured to provide location information (e.g., longitude and latitude) about the current location of the WTRU 102. In addition to, or instead of, information from the GPS chipset 136, the WTRU 102 may receive location information from base stations (e.g., base stations 114a, 114b) via air interface 116 and / or determine its location based on the timing of signals received from two or more nearby base stations. It will be understood that the WTRU 102 may acquire location information using any suitable location determination method, while remaining consistent with the embodiments.

[0060] The processor 118 may be further coupled to other peripherals 138, which may include one or more software and / or hardware modules providing additional features, functions, and / or wired or wireless connectivity. For example, peripherals 138 may include accelerometers, electronic compasses, satellite transceivers, digital cameras (for photos and / or videos), Universal Serial Bus (USB) ports, vibration devices, television transceivers, hands-free headsets, Bluetooth® modules, FM radio units, digital music players, media players, video game player modules, internet browsers, virtual reality and / or augmented reality (VR / AR) devices, activity trackers, etc. Peripherals 138 may include one or more sensors, which may be one or more of the following: gyroscopes, accelerometers, Hall effect sensors, magnetometers, orientation sensors, proximity sensors, temperature sensors, time sensors, geolocation sensors, altimeters, light sensors, touch sensors, magnetometers, barometers, gesture sensors, biometric sensors, and / or humidity sensors.

[0061] WTRU 102 may include a full-duplex radio, for which the transmission and reception of some or all signals (e.g., associated with specific subframes for both UL (e.g., for transmission) and downlink (e.g., for reception)) may be concurrent and / or simultaneous. The full-duplex radio may include an interference management unit to reduce and / or substantially eliminate self-interference via hardware (e.g., a choke) or via signal processing (e.g., a separate processor (not shown) or via processor 118). In one embodiment, WTRU 102 may include a half-duplex radio, for which the transmission and reception of some or all signals (e.g., associated with specific subframes for UL (e.g., for transmission) or downlink (e.g., for reception)) may be concurrent and / or simultaneous.

[0062] Figure 1C The diagram illustrates a system diagram of RAN 104 and CN 106 according to an embodiment. As noted above, RAN 104 can employ E-UTRA radio technology to communicate with WTRUs 102a, 102b, and 102c via air interface 116. RAN 104 can also communicate with CN 106.

[0063] RAN 104 may include eNode-Bs 160a, 160b, and 160c, although it will be understood that RAN 104 may include any number of eNode-Bs while remaining consistent with the embodiments. eNode-Bs 160a, 160b, and 160c may each include one or more transceivers for communicating with WTRUs 102a, 102b, and 102c via air interface 116. In one embodiment, eNode-Bs 160a, 160b, and 160c may implement MIMO technology. Therefore, eNode-B 160a may, for example, use multiple antennas to transmit radio signals to and / or receive radio signals from WTRU 102a.

[0064] Each of the eNode-B 160a, 160b, and 160c can be associated with a specific cell (not shown) and can be configured to handle radio resource management decisions, handover decisions, and user scheduling in the UL and / or DL, etc. Figure 1C As shown, eNode-B 160a, 160b, and 160c can communicate with each other via the X2 interface.

[0065] Figure 1C The CN 106 shown may include a Mobility Management Entity (MME) 162, a Serving Gateway (SGW) 164, and a Packet Data Network (PDN) Gateway (or PGW) 166. Although each of the foregoing elements is depicted as part of CN 106, it will be understood that any of these elements may be owned and / or operated by an entity other than a CN operator.

[0066] The MME 162 can connect to each of the eNode-Bs 162a, 162b, and 162c in RAN104 via the S1 interface and can act as a control node. For example, the MME 162 can be responsible for authenticating users of WTRUs 102a, 102b, and 102c, activating / deactivating bearers, and selecting a specific serving gateway during the initial attachment of WTRUs 102a, 102b, and 102c. The MME 162 can provide control plane functions for handover between RAN104 and other RANs (not shown) employing other radio technologies, such as GSM and / or WCDMA.

[0067] The SGW 164 can connect to each of the eNode Bs 160a, 160b, and 160c in RAN104 via the S1 interface. The SGW 164 can typically route and forward user data packets to / from WTRUs 102a, 102b, and 102c. The SGW 164 can perform other functions, such as anchoring the user plane during inter-eNode B handover, triggering paging when DL data is available for WTRUs 102a, 102b, and 102c, and managing and storing the context of WTRUs 102a, 102b, and 102c.

[0068] SGW 164 can be connected to PGW 166, which can provide WTRU 102a, 102b, 102c with access to packet-switched networks (such as Internet 110) to facilitate communication between WTRU 102a, 102b, 102c and IP-enabled devices.

[0069] CN 106 facilitates communication with other networks. For example, CN 106 can provide WTRUs 102a, 102b, and 102c with access to a circuit-switched network (such as PSTN 108) to facilitate communication between WTRUs 102a, 102b, and 102c and conventional terrestrial line communication equipment. For example, CN 106 may include or communicate with an IP gateway (e.g., an IP Multimedia Subsystem (IMS) server), which acts as an interface between CN 106 and PSTN 108. Additionally, CN 106 can provide WTRUs 102a, 102b, and 102c with access to other networks 112, which may include other wired and / or wireless networks owned and / or operated by other service providers.

[0070] Despite WTRU in Figures 1A to 1D While described as a wireless terminal, it is envisioned that, in some representative embodiments, such a terminal may use (e.g., temporarily or permanently) a wired communication interface with a communication network.

[0071] In a representative embodiment, the other network 112 may be a WLAN.

[0072] In an Infrastructure Basic Services Set (BSS) mode, a WLAN may have an access point (AP) for the BSS and one or more stations (STAs) associated with the AP. The AP may have access or an interface to a distribution system (DS) or another type of wired / wireless network that carries traffic into and / or out of the BSS. Traffic originating outside the BSS destined for a STA can be delivered to the STA via the AP. Traffic from a STA to a destination outside the BSS can be sent to the AP for delivery to the appropriate destination. Traffic between STAs within the BSS can be sent via the AP, for example, where a source STA can send traffic to the AP, and the AP can deliver the traffic to the destination STA. Traffic between STAs within the BSS can be considered and / or referred to as peer-to-peer traffic. Peer-to-peer traffic can be sent between source and destination STAs (e.g., directly between them) using a direct link setup (DLS). In some representative embodiments, the DLS may use 802.11e DLS or 802.11z Tunneled DLS (TDLS). A WLAN using the Standalone BSS (IBSS) mode may not have an access point (AP), and STAs within or using the IBSS (e.g., all STAs) can communicate directly with each other. The IBSS communication mode may sometimes be referred to as a "self-organizing" communication mode in this document.

[0073] When using 802.11ac infrastructure operation mode or a similar operation mode, the AP can transmit beacons on a fixed channel, such as a primary channel. The primary channel can be of fixed width (e.g., a bandwidth of 20 MHz) or dynamically set via signaling. The primary channel can be the operating channel of the BSS and can be used by STAs to establish connections with the AP. In some representative embodiments, Carrier Sense Multiple Access with Collision Avoidance (CSMA / CA) can be implemented, for example, in an 802.11 system. For CSMA / CA, STAs including the AP (e.g., each STA) can sense the primary channel. If the primary channel is sensed / detected and / or determined to be busy by a particular STA, that STA can back off. A STA (e.g., only one station) can transmit at any given time within a given BSS.

[0074] High-throughput (HT) STAs can communicate using a 40MHz wide channel, for example, by combining a primary 20MHz channel with adjacent or non-adjacent 20MHz channels to form a 40MHz wide channel.

[0075] Very High Throughput (VHT) STAs can support channels with widths of 20MHz, 40MHz, 80MHz, and / or 160MHz. 40MHz and / or 80MHz channels can be formed by combining consecutive 20MHz channels. A 160MHz channel can be formed by combining eight consecutive 20MHz channels or by combining two non-consecutive 80MHz channels (which can be referred to as an 80+80 configuration). For the 80+80 configuration, after channel coding, data is transmitted via a segment resolver that divides the data into two streams. Inverse Fast Fourier Transform (IFFT) processing and time-domain processing can be performed separately on each stream. The streams can be mapped onto two 80MHz channels, and the data can be transmitted by the transmitting STA. At the receiver of the receiving STA, the above operations for the 80+80 configuration can be reversed, and the combined data can be sent to the Media Access Control (MAC).

[0076] Operating modes below 1 GHz are supported by 802.11af and 802.11ah. The channel operating bandwidth and carrier are reduced in 802.11af and 802.11ah compared to those used in 802.11n and 802.11ac. 802.11af supports 5 MHz, 10 MHz, and 20 MHz bandwidths in the TV white space (TVWS) spectrum, while 802.11ah supports 1 MHz, 2 MHz, 4 MHz, 8 MHz, and 16 MHz bandwidths using non-TVWS spectrum. According to representative embodiments, 802.11ah can support instrument-type control / machine-type communications, such as MTC devices in macro coverage areas. MTC devices may have certain capabilities, such as limited capabilities, including support for (e.g., only support) certain and / or limited bandwidths. MTC devices may include batteries with a lifespan exceeding a threshold (e.g., to maintain a very long battery life).

[0077] WLAN systems that can support multiple channels and channel bandwidths (such as 802.11n, 802.11ac, 802.11af, and 802.11ah) include a channel that can be designated as the primary channel. The primary channel can have a bandwidth equal to the maximum common operating bandwidth supported by all STAs in the BSS. The bandwidth of the primary channel can be set and / or limited by the STA that supports the minimum bandwidth operating mode among all STAs operating in the BSS. In the 802.11ah example, for STAs that support (e.g., only support) the 1MHz mode (e.g., MTC type devices), the primary channel can be 1MHz wide, even if the AP and other STAs in the BSS support 2MHz, 4MHz, 8MHz, 16MHz, and / or other channel bandwidth operating modes. Carrier sensing and / or Network Allocation Vector (NAV) settings can depend on the status of the primary channel. If the primary channel is busy, for example because an STA (which only supports the 1MHz operating mode) is transmitting to the AP, the entire available band may be considered busy, even if most of the band is still idle and could be available.

[0078] In the United States, the available frequency band for 802.11ah is from 902MHz to 928MHz. In South Korea, the available frequency band is from 917.5MHz to 923.5MHz. In Japan, the available frequency band is from 916.5MHz to 927.5MHz. Depending on the country code, the total available bandwidth for 802.11ah is 6MHz to 26MHz.

[0079] Figure 1D The diagram illustrates a system diagram of RAN 113 and CN 115 according to an embodiment. As noted above, RAN 113 may employ NR radio technology to communicate with WTRUs 102a, 102b, and 102c via air interface 116. RAN 113 may also communicate with CN 115.

[0080] RAN 113 may include gNBs 180a, 180b, and 180c, although it will be understood that RAN 113 may include any number of gNBs while remaining consistent with the embodiments. gNBs 180a, 180b, and 180c may each include one or more transceivers for communicating with WTRUs 102a, 102b, and 102c via air interface 116. In one embodiment, gNBs 180a, 180b, and 180c may implement MIMO technology. For example, gNBs 180a and 180b may utilize beamforming to transmit signals to and / or receive signals from gNBs 180a, 180b, and 180c. Therefore, gNB 180a may, for example, use multiple antennas to transmit radio signals to and / or receive radio signals from WTRU 102a. In one embodiment, gNBs 180a, 180b, and 180c can implement carrier aggregation technology. For example, gNB 180a can transmit multiple component carriers to WTRU 102a (not shown). A subset of these component carriers can be on unlicensed spectrum, while the remaining component carriers can be on licensed spectrum. In one embodiment, gNBs 180a, 180b, and 180c can implement Cooperative Multipoint (CoMP) technology. For example, WTRU 102a can receive cooperative transmissions from gNBs 180a and 180b (and / or gNB 180c).

[0081] WTRUs 102a, 102b, and 102c can communicate with gNBs 180a, 180b, and 180c using transmissions associated with scalable digitization. For example, OFDM symbol spacing and / or OFDM subcarrier spacing can vary depending on different transmissions, different cells, and / or different portions of the radio transmission spectrum. WTRUs 102a, 102b, and 102c can communicate with gNBs 180a, 180b, and 180c using subframes of various or scalable lengths or transmission time intervals (TTIs) (e.g., containing different numbers of OFDM symbols and / or absolute times of varying durations).

[0082] gNBs 180a, 180b, and 180c can be configured to communicate with WTRUs 102a, 102b, and 102c in standalone and / or non-standalone configurations. In standalone configuration, WTRUs 102a, 102b, and 102c can communicate with gNBs 180a, 180b, and 180c without also accessing other RANs (e.g., eNode-Bs 160a, 160b, and 160c). In standalone configuration, WTRUs 102a, 102b, and 102c can utilize one or more of gNBs 180a, 180b, and 180c as mobility anchors. In standalone configuration, WTRUs 102a, 102b, and 102c can communicate with gNBs 180a, 180b, and 180c using signals in unlicensed frequency bands. In a non-standalone configuration, WTRUs 102a, 102b, and 102c can communicate with / connect to gNBs 180a, 180b, and 180c, and simultaneously communicate with / connect to another RAN (such as eNode-Bs 160a, 160b, and 160c). For example, WTRUs 102a, 102b, and 102c can implement DC principles to communicate substantially simultaneously with one or more gNBs 180a, 180b, and 180c and one or more eNode-Bs 160a, 160b, and 160c. In a non-standalone configuration, eNode-B 160a, 160b, and 160c can act as mobility anchors for WTRU 102a, 102b, and 102c, and gNB 180a, 180b, and 180c can provide additional coverage and / or throughput to serve WTRU 102a, 102b, and 102c.

[0083] Each of gNBs 180a, 180b, and 180c can be associated with a specific cell (not shown) and can be configured to handle radio resource management decisions, handover decisions, user scheduling in UL and / or DL, network slicing support, dual connectivity, interoperability between NR and E-UTRA, routing of user plane data to User Plane Functions (UPF) 184a and 184b, routing of control plane information to Access and Mobility Management Functions (AMF) 182a and 182b, etc. Figure 1D As shown, gNB 180a, 180b, and 180c can communicate with each other via the Xn interface.

[0084] Figure 1DThe CN 115 shown may include at least one AMF 182a, 182b, at least one UPF 184a, 184b, at least one Session Management Function (SMF) 183a, 183b, and possibly a Data Network (DN) 185a, 185b. Although each of the foregoing elements is depicted as part of the CN 115, it will be understood that any of these elements may be owned and / or operated by an entity other than the CN operator.

[0085] AMF 182a and 182b can connect to one or more of the gNBs 180a, 180b, and 180c in RAN 113 via the N2 interface and can act as control nodes. For example, AMF 182a and 182b can be responsible for authenticating users of WTRU 102a, 102b, and 102c, supporting network slicing (e.g., handling different PDU sessions with different requirements), selecting specific SMF183a and 183b, managing registration areas, terminating NAS signaling, mobility management, etc. Network slices can be used by AMF182a and 182b to customize CN support for WTRU 102a, 102b, and 102c based on the service types utilized by WTRU 102a, 102b, and 102c. For example, different network slices can be built for different use cases, such as services that rely on Ultra Reliable Low Latency (URLLC) access, services that rely on Enhanced Massive Mobile Broadband (eMBB) access, and services for Machine Type Communication (MTC) access. AMF 162 can provide control plane functions for handover between RAN 113 and other RANs (not shown) employing other radio technologies, such as LTE, LTE-A, LTE-A Pro, and / or non-3GPP access technologies, such as WiFi.

[0086] SMFs 183a and 183b can connect to AMFs 182a and 182b in CN 115 via the N11 interface. SMFs 183a and 183b can also connect to UPFs 184a and 184b in CN 115 via the N4 interface. SMFs 183a and 183b can select and control UPFs 184a and 184b, and configure traffic routing through UPFs 184a and 184b. SMFs 183a and 183b can perform other functions, such as managing and allocating UE IP addresses, managing PDU sessions, controlling policy enforcement and QoS, and providing downlink data notifications. PDU session types can be IP-based, non-IP-based, Ethernet-based, etc.

[0087] UPF 184a and 184b can connect to one or more of gNBs 180a, 180b, and 180c in RAN 113 via the N3 interface. These gNBs can provide WTRU 102a, 102b, and 102c with access to a packet-switched network (such as the Internet 110) to facilitate communication between WTRU 102a, 102b, 102c and IP-enabled devices. UPF 184a and 184b can perform other functions such as routing and forwarding packets, enforcing user plane policies, supporting multi-homed PDU sessions, handling user plane QoS, buffering downlink packets, and providing mobility anchoring.

[0088] CN 115 can facilitate communication with other networks. For example, CN 115 may include or communicate with an IP gateway (e.g., an IP Multimedia Subsystem (IMS) server) that acts as an interface between CN 115 and PSTN 108. Additionally, CN 115 can provide WTRUs 102a, 102b, and 102c with access to other networks 112, which may include other wired and / or wireless networks owned and / or operated by other service providers. In one embodiment, WTRUs 102a, 102b, and 102c can connect to local data networks (DNs) 185a and 185b via the N3 interface to UPFs 184a and 184b and the N6 interface between UPFs 184a and 184b and DNs 185a and 185b.

[0089] Given Figures 1A to 1D as well as Figures 1A to 1D The corresponding descriptions may be performed by one or more emulation devices (not shown) to perform one or more of the functions described herein with respect to one or more of the following: WTRU102a-d, base station 114a-b, eNode-B160a-c, MME 162, SGW 164, PGW 166, gNB180a-c, AMF 182a-b, UPF 184a-b, SMF 183a-b, DN185a-b, and / or one or more other devices described herein. An emulation device may be one or more devices configured to emulate one or more of the functions described herein. For example, an emulation device may be used to test other devices and / or simulate network and / or WTRU functions.

[0090] Simulation devices can be designed to perform tests on one or more other devices in a laboratory environment and / or a carrier network environment. For example, the one or more simulation devices can perform one or more or all of their functions while being fully or partially implemented and / or deployed as part of a wired and / or wireless communication network to test other devices within the communication network. The one or more simulation devices can perform one or more or all of their functions while being temporarily implemented / deployed as part of a wired and / or wireless communication network. Simulation devices can be directly coupled to another device for testing purposes and / or can use over-the-air wireless communication to perform tests.

[0091] The one or more simulation devices can perform one or more (including all) functions without being implemented / deployed as part of a wired and / or wireless communication network. For example, the simulation devices can be used in test scenarios in a test laboratory and / or in non-deployed (e.g., testing) wired and / or wireless communication networks to perform testing on one or more components. The one or more simulation devices can be test rigs. Direct RF coupling and / or wireless communication via RF circuitry (e.g., which may include one or more antennas) can be used by the simulation devices to transmit and / or receive data.

[0092] This application describes various aspects, including tools, features, examples, models, schemes, etc. Many of these aspects are described in detail and are often described in a way that may sound restrictive, at least to illustrate individual characteristics. However, this is for clarity of purpose and does not limit the application or scope of those aspects. Indeed, all the different aspects can be combined and interchanged to provide further aspects. Furthermore, this aspect can also be combined and interchanged with aspects described in earlier filings.

[0093] The aspects described and conceived in this application can be implemented in many different forms. Figure 5-12 Some examples can be provided, but other examples can be thought of. Figure 5-12 The discussion does not limit the breadth of implementation methods. At least one aspect generally relates to video encoding and decoding, and at least one other aspect generally relates to the transmission of a generated or encoded bitstream. These and other aspects can be implemented as methods, apparatus, computer-readable storage media having instructions stored thereon for encoding or decoding video data according to any of the described methods, and / or computer-readable storage media having bitstreams generated according to any of the described methods stored thereon.

[0094] In this application, the terms “reconstructed” and “decoded” may be used interchangeably, as may the terms “pixel” and “sample”, and may the terms “image”, “picture” and “frame” be used interchangeably.

[0095] This document describes various methods, each of which includes one or more steps or actions for implementing the described method. Unless a specific order of steps or actions is required for the proper operation of the method, the order and / or use of specific steps and / or actions may be modified or combined. Additionally, terms such as "first," "second," etc., may be used in various examples to modify elements, components, steps, operations, etc., such as, for example, "first decoding" and "second decoding." The use of such terms does not imply a sequence of operations unless specifically required. Thus, in this example, the first decoding need not be performed before the second decoding, but may occur, for example, before, during, or in the time period overlapping with the second decoding.

[0096] The various methods and other aspects described in this application can be used to modify, for example Figure 2 and Figure 3 The modules of the video encoder 200 and decoder 300 shown herein, for example, a decoding module. Furthermore, the subject matter disclosed herein can be applied, for example, to any type, format, or version of video encoding, whether or not described in a standard or recommendation, whether pre-existing or future-developed, and any extensions to such standards and recommendations. Unless otherwise indicated or technically excluded, the aspects described herein may be used individually or in combination.

[0097] Various numerical values, such as 1, 2, 4, 7, 8, 16, 32, 64, etc., are used in the examples described in this application. These and other specific values ​​are used for the purpose of describing the examples, and the aspects described are not limited to these specific values.

[0098] Figure 2 This is a diagram illustrating an example video encoder 200. Variations of the example encoder 200 can be conceived, but for clarity, the encoder 200 is described below without depicting all expected variations.

[0099] Before being encoded, the video sequence may undergo pre-encoding processing (201), such as applying color transformations to the input color image (e.g., a conversion from RGB 4:4:4 to YCbCr 4:2:0), or performing remapping of the input image components to obtain a more resilient signal distribution for compression (e.g., histogram equalization using one of the color components). Metadata (e.g., which may include film grain parameters determined through preprocessing as described herein) may be associated with the preprocessing and attached to the bitstream.

[0100] In encoder 200, the image is encoded by encoder elements as described below. The image to be encoded is partitioned (202) and processed in units such as coding units (CUs). Each unit is encoded using, for example, an intra-frame or inter-frame mode. When a unit is encoded in an intra-frame mode, it performs intra-frame prediction (260). In an inter-frame mode, motion estimation (275) and compensation (270) are performed. The encoder determines (205) which of the intra-frame or inter-frame modes to use for encoding the unit and indicates the intra-frame / inter-frame decision by, for example, a prediction mode flag. The prediction residual is calculated, for example, by subtracting (210) the prediction block from the original image block.

[0101] The predicted residual is then transformed (225) and quantized (230). The quantized transform coefficients, motion vectors, and other syntax elements are entropy encoded (245) to output a bit stream. The encoder can skip the transform and apply the quantization directly to the untransformed residual signal. The encoder can bypass both the transform and quantization, i.e., encode the residual directly without applying the transform or quantization process.

[0102] The encoder decodes the encoded blocks to provide a reference for further prediction. The quantized transform coefficients are dequantized (240) and inverse transformed (250) to decode the prediction residuals. The decoded prediction residuals and the predicted blocks are combined (255) to reconstruct the image blocks. A loop filter (265) is applied to the reconstructed image to perform, for example, deblocking / SAO (sample adaptive offset) filtering to reduce coding artifacts. The filtered image is stored in a reference image buffer (280).

[0103] Figure 3 This is a diagram illustrating an example video decoder. In the example decoder 300, the bitstream is decoded by decoder elements, as described below. The video decoder 300 generally performs the same operations as... Figure 2 The encoding passes described herein are mutually decoded passes. Encoder 200 generally also performs video decoding as part of the encoding of video data.

[0104] Specifically, the input to the decoder includes a video bitstream, which can be generated by the video encoder 200. First, entropy decoding (330) is performed on the bitstream to obtain transform coefficients, motion vectors, and other encoded information. Image partitioning information indicates how the image should be partitioned. Therefore, the decoder can partition (335) the image based on the decoded image partitioning information. The transform coefficients are dequantized (340) and inverse transformed (350) to decode the prediction residuals. The decoded prediction residuals and the predicted blocks are combined (355) to reconstruct the image blocks. The predicted blocks (370) can be obtained from intra-frame prediction (360) or motion-compensated prediction (i.e., inter-frame prediction) (375). A loop filter (365) is applied to the reconstructed image. The filtered image is stored in a reference image buffer (380).

[0105] The decoded image can undergo further post-decoding processing (385), such as inverse color transformation (e.g., conversion from YCbCr 4:2:0 to RGB 4:4:4), or inverse remapping, performing the inverse operation of the remapping process performed in pre-encoding processing (201). Post-decoding processing can utilize metadata derived in pre-encoding processing and signaled in the bitstream. In one example, the decoded image (e.g., after the application of a loop filter (365) and / or after post-decoding processing (385) in the case of post-decoding processing) can be sent to a display device for presentation to a user.

[0106] Figure 4 This is a diagram illustrating examples of systems in which the various aspects and examples described herein can be implemented. System 400 may be embodied as a device including the various components described below and configured to perform one or more of the aspects described in this document. Examples of such devices include, but are not limited to, various electronic devices such as personal computers, laptop computers, smartphones, tablet computers, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected home appliances, and servers. Elements of system 400 may be embodied individually or in combination in a single integrated circuit (IC), multiple ICs, and / or discrete components. For example, in at least one example, the processing and encoder / decoder elements of system 400 are distributed across multiple ICs and / or discrete components. In various examples, system 400 is communicatively coupled to one or more other systems or other electronic devices via, for example, a communication bus or through dedicated input and / or output ports. In various examples, system 400 is configured to implement one or more of the aspects described in this document.

[0107] System 400 includes: at least one processor 410 configured to execute instructions loaded therein for implementing various aspects, such as those described in this document. Processor 410 may include embedded memory, input / output interfaces, and various other circuitry as known in the art. System 400 includes at least one memory 420 (e.g., a volatile memory device and / or a non-volatile memory device). System 400 includes: a storage device 440 which may include non-volatile memory and / or volatile memory, including but not limited to electrically erasable programmable read-only memory (EEPROM), read-only memory (ROM), programmable read-only memory (PROM), random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), flash memory, disk drives, and / or optical disk drives. Storage device 440 may include internal storage devices, attached storage devices (including removable and non-removable storage devices), and / or network-accessible storage devices, as non-limiting examples.

[0108] System 400 includes an encoder / decoder module 430 configured to, for example, process data to provide encoded or decoded video, and the encoder / decoder module 430 may include its own processor and memory. The encoder / decoder module 430 represents one or more modules that may be included in a device to perform encoding and / or decoding functions. It is well known that a device may include one or both encoding and decoding modules. Furthermore, the encoder / decoder module 430 may be implemented as a separate element of system 400, or may be incorporated into processor 410 as a combination of hardware and software as known to those skilled in the art.

[0109] Program code to be loaded onto processor 410 or encoder / decoder 430 to execute the various aspects described herein may be stored in storage device 440 and subsequently loaded onto memory 420 for execution by processor 410. According to various examples, one or more of processor 410, memory 420, storage device 440, and encoder / decoder module 430 may store one or more items of various kinds during the execution of the processes described herein. Such stored items may include, but are not limited to, input video, decoded video or portions of decoded video, bitstreams, matrices, variables, and intermediate or final results from the processing of equations, formulas, operations, and operational logic.

[0110] In some examples, the memory within processor 410 and / or encoder / decoder module 430 is used to store instructions and provide working memory for processing required during encoding or decoding. However, in other examples, external memory (e.g., processor 410 or encoder / decoder module 430) is used for one or more of these functions. External memory may be memory 420 and / or storage device 440, such as volatile memory and / or non-volatile flash memory. In several examples, external non-volatile flash memory is used to store, for example, the operating system of a television. In at least one example, fast external volatile memory, such as RAM, is used as working memory for video encoding and decoding operations.

[0111] Input to the components of system 400 can be provided through various input devices as indicated in box 445. Such input devices include, but are not limited to: (i) a radio frequency (RF) section that receives RF signals transmitted over the air, for example, by a broadcaster; (ii) component (COMP) input terminals (or a collection of COMP input terminals); (iii) a universal serial bus (USB) input terminal; and / or (iv) a high-definition multimedia interface (HDMI) input terminal. Figure 4 Other examples not shown include composite video.

[0112] In various examples, the input device of block 445 has associated corresponding input processing elements as known in the art. For example, the RF section may be associated with elements suitable for: (i) selecting a desired frequency (also referred to as selecting a signal or limiting a signal to a frequency band); (ii) down-converting the selected signal; (iii) further limiting the frequency band to a narrower band to select a signal band that may be referred to as a channel in some examples; (iv) demodulating the down-converted and band-limited signal; (v) performing error correction; and / or (vi) demultiplexing to select the desired stream of data packets. The RF section of various examples includes one or more elements to perform these functions, such as frequency selectors, signal selectors, band limiters, channel selectors, filters, downconverters, demodulators, error correctors, and demultiplexers. The RF section may include: a tuner that performs various of these functions, including, for example, down-converting a received signal to a lower frequency (e.g., intermediate frequency or near-baseband frequency) or to baseband. In one set-top box example, the RF section and its associated input processing elements receive RF signals transmitted over a wired (e.g., cable) medium and perform frequency selection by filtering, down-converting, and filtering again to the desired frequency band. Various examples rearrange the order of the components described above (and others), remove some of these components, and / or add other components that perform similar or different functions. Adding components may include inserting components between existing components, such as, for example, inserting amplifiers and analog-to-digital converters. In various examples, the RF section includes an antenna.

[0113] USB and / or HDMI endpoints may include corresponding interface processors for connecting system 400 to other electronic devices across USB and / or HDMI connections. It should be understood that various aspects of input processing, such as Reed-Solomon error correction, may be implemented, for example, within a separate input processing IC or, if necessary, within processor 410. Similarly, aspects of USB or HDMI interface processing may be implemented, either within a separate interface IC or, if necessary, within processor 410. The demodulated, error-corrected, and demultiplexed stream is provided to various processing elements, including, for example, processor 410 and encoder / decoder 430, which operate in conjunction with memory and storage elements to process the data stream as needed for presentation on an output device.

[0114] Various components of the system 400 can be provided within an integrated housing, in which the various components can be interconnected and data can be transmitted therebetween using a suitable connection arrangement 425, such as an internal bus as known in the art, including inter-IC (I2C) bus, wiring and printed circuit board.

[0115] System 400 includes a communication interface 450, which enables communication with other devices via a communication channel 460. The communication interface 450 may include, but is not limited to, a transceiver configured to transmit and receive data on the communication channel 460. The communication interface 450 may include, but is not limited to, a modem or network interface card (NIC), and the communication channel 460 may be implemented over, for example, wired and / or wireless media.

[0116] In various examples, data is streamed or otherwise provided to system 400 using a wireless network such as WiFi (e.g., IEEE 802.11 (IEEE refers to the Institute of Electrical and Electronics Engineers)). In these examples, the Wi-Fi signal is received on a communication channel 460 and a communication interface 450 adapted for Wi-Fi communication. The communication channel 460 in these examples is typically connected to an access point or router that provides access to external networks, including the Internet, to allow streaming applications and other over-the-top communications. Other examples use a set-top box that delivers data over an HDMI connection in input box 445 to provide streaming data to system 400. Still other examples use an RF connection in input box 445 to provide streaming data to system 400. As indicated above, various examples provide data in a non-streaming manner. Additionally, various examples use wireless networks other than Wi-Fi, such as cellular networks or Bluetooth® networks.

[0117] System 400 can provide output signals to various output devices, including a display 475, a speaker 485, and other peripheral devices 495. Various examples of the display 475 include one or more of the following: for example, a touchscreen display, an organic light-emitting diode (OLED) display, a curved display, and / or a foldable display. The display 475 can be used in a television, tablet, laptop, cellular phone, or another device. The display 475 can also be integrated with other components (e.g., as in a smartphone) or separate (e.g., an external monitor for a laptop). In various examples, other peripheral devices 495 include one or more of a standalone digital video disc (or digital multifunction disc) (DVD, for both terms), a disc player, a stereo system, and / or a lighting system. Various examples utilize one or more peripheral devices 495 that provide functionality based on the output of system 400. For example, a disc player performs the function of playing the output of system 400.

[0118] In various examples, signaling such as AV is used to transmit control signals between system 400 and display 475, speaker 485, or other peripheral devices 495. Device-to-device control links, consumer electronics control (CEC), or other communication protocols are implemented with or without user intervention. Output devices can be communicatively coupled to system 400 via dedicated connections through corresponding interfaces 470, 480, and 490. Alternatively, output devices can be connected to system 400 via communication interface 450 using communication channel 460. Display 475 and speaker 485 can be integrated into a single unit with other components of system 400 in electronic devices such as, for example, televisions. In various examples, display interface 470 includes display drivers, such as, for example, timing controller (TCon) chips.

[0119] Display 475 and speaker 485 can alternatively be separated from one or more other components, for example, if the RF section of input 445 is part of a separate set-top box. In various examples where display 475 and speaker 485 are external components, output signals can be provided via dedicated output connections, including, for example, HDMI ports, USB ports, or COMP outputs.

[0120] The example can be implemented by computer software, hardware, or a combination of hardware and software, as implemented by processor 410. As a non-limiting example, the example can be implemented by one or more integrated circuits. As a non-limiting example, memory 420 can be of any type suitable for the technical environment and can be implemented using any suitable data storage technology, such as optical memory devices, magnetic memory devices, semiconductor-based memory devices, fixed memory, and removable memory. As a non-limiting example, processor 410 can be of any type suitable for the technical environment and can encompass one or more of microprocessors, general-purpose computers, special-purpose computers, and processors based on multi-core architectures.

[0121] Various implementations involve decoding. As used in this application, "decoding" can encompass all or part of a process performed, for example, on a received encoded sequence to produce a final output suitable for display. In various examples, such a process includes one or more processes typically performed by a decoder, such as entropy decoding, dequantization, inverse transform, and differential decoding. In various examples, this process may also alternatively include processes performed by decoders of various implementations described in this application, such as: determining a decoder-side intra-mode derivation (DIMD) merging mode enabled for the current block; determining reference samples to be used for the DIMD merging mode, wherein the reference samples include a first reference sample associated with a first gradient histogram and a second reference sample associated with a second gradient histogram; scaling the first gradient histogram based on a first distance between the first reference sample and the current block; scaling the second gradient histogram based on a second distance between the second reference sample and the current block; deriving a merged gradient histogram based on the scaled first gradient histogram and the scaled second gradient histogram; deriving one or more intra-prediction modes based on the merged gradient histogram; obtaining a prediction for the current block based on the one or more intra-prediction modes; and decoding the current block based on the prediction.

[0122] As further examples, in one example, "decoding" refers only to entropy decoding; in another example, "decoding" refers only to differential decoding; and in yet another example, "decoding" refers to a combination of entropy decoding and differential decoding. Whether the phrase "decoding process" is intended to specifically refer to a subset of operations or generally to a broader decoding process will be clear based on the specific context of the description and is considered to be well understood by those skilled in the art.

[0123] Various implementations involve encoding. In a manner similar to the discussion above regarding “decoding,” the term “encoding,” as used herein, can encompass all or part of a process performed, for example, on an input video sequence to produce an encoded bitstream. In various examples, this process includes one or more processes typically performed by an encoder, such as partitioning, differential coding, transform, quantization, and entropy coding. In various examples, this process may also alternatively include processes performed by encoders of various implementations described in this application, such as: determining a decoder-side intra-frame mode derivation (DIMD) merging mode enabled for the current block; determining reference samples to be used for the DIMD merging mode, wherein the reference samples include a first reference sample associated with a first gradient histogram and a second reference sample associated with a second gradient histogram; scaling the first gradient histogram based on a first distance between the first reference sample and the current block; scaling the second gradient histogram based on a second distance between the second reference sample and the current block; deriving a merged gradient histogram based on the scaled first gradient histogram and the scaled second gradient histogram; and encoding the current block based on the merged gradient histogram.

[0124] As further examples, in one example, "encoding" refers only to entropy encoding; in another example, "encoding" refers only to differential encoding; and in yet another example, "encoding" refers to a combination of differential and entropy encoding. Whether the phrase "encoding process" is intended to specifically refer to a subset of operations or generally to a broader encoding process will be clear based on the specific context of the description and is considered well understood by those skilled in the art.

[0125] Note that the syntax elements used in this article (e.g., coding syntax on intensity intervals, granular parameters, block offsets, expansion factors, etc.) are descriptive terms. Therefore, they do not preclude the use of other syntax element names.

[0126] When a diagram is presented as a flowchart, it should be understood that it also provides a block diagram of the corresponding device. Similarly, when a diagram is presented as a block diagram, it should be understood that it also provides a flowchart of the corresponding method / process.

[0127] The implementations and aspects described herein can be implemented, for example, in methods or processes, apparatus, software programs, data streams, or signals. Even if discussed only in the context of a single form of implementation (e.g., discussed only as a method), the features in question can be implemented in other forms (e.g., apparatus or program). Apparatus can be implemented, for example, in appropriate hardware, software, and firmware. Methods can be implemented, for example, in a processor, where processor generally refers to a processing device, including, for example, a computer, microprocessor, integrated circuit, or programmable logic device. Processors also include communication devices, such as, for example, computers, cellular phones, portable / personal digital assistants (“PDAs”), and other devices that facilitate communication of information between end users.

[0128] References to “an example” or “an instance” or “an implementation” or “an implementation” and their variations mean that the specific feature, structure, characteristic, etc., described in connection with the example is included in at least one example. Therefore, the appearance of the phrase “in an example” or “in an instance” or “in an implementation” or “in an implementation” and any variations appearing throughout this application do not necessarily all refer to the same example.

[0129] Additionally, this application may refer to "determining" various pieces of information. Determining information may include one or more of the following: estimated information, calculated information, predicted information, or information retrieved from memory. Obtaining may include receiving, retrieving, constructing, generating, and / or determining.

[0130] Furthermore, this application may refer to "accessing" various pieces of information. Accessing information may include one or more of the following: receiving information, retrieving information (e.g., from memory), storing information, moving information, copying information, calculating information, determining information, predicting information, or estimating information.

[0131] Additionally, this application may refer to "receiving" various pieces of information. As with "access," "receiving" is intended to be a broad term. Receiving information may include one or more of, for example, accessing information or retrieving information (e.g., from memory). Further, "receiving" typically refers to actions performed during operation, such as storing information, processing information, transmitting information, moving information, copying information, erasing information, calculating information, determining information, predicting information, or estimating information.

[0132] It should be understood that the use of any of the following “ / ”, “and / or”, and “…at least one of” (e.g., in the cases of “A / B”, “A and / or B”, and “at least one of A and B”) is intended to cover the selection of only the first listed option (A), or only the second listed option (B), or the selection of both options (A and B). As a further example, in the cases of “A, B, and / or C” and “at least one of A, B, and C”, this phrase is intended to cover the selection of only the first listed option (A), or only the second listed option (B), or only the third listed option (C), or only the first and second listed options (A and B), or only the first and third listed options (A and C), or only the second and third listed options (B and C), or the selection of all three options (A, B, and C). This can be extended to as many items as are listed, as will be clear to those skilled in the art and related fields.

[0133] Moreover, as used herein, the term "signaling" refers, among other things, to instructing the corresponding decoder to do something. Encoder signals may include, for example, the number of intensity intervals, the number of model values, particle parameters, particle identifiers, scaling factors, and so on. In this way, in one example, the same parameters are used on both the encoder and decoder sides. Thus, for example, the encoder can transmit specific parameters (explicit signaling) to the decoder so that the decoder can use the same specific parameters. Conversely, if the decoder already has specific parameters as well as other parameters, then signaling can be used without transmission (implicit signaling) to simply allow the decoder to know and select specific parameters. Bit saving is achieved in various examples by avoiding the transmission of any actual functionality. It should be understood that signaling can be done in a variety of ways. For example, in various examples, one or more syntax elements, tags, etc., are used to signal information to the corresponding decoder. Although the foregoing refers to the verb form of the term "signaling," the term "signaling" may also be used as a noun in this text.

[0134] As will be apparent to those skilled in the art, implementations can generate various signals formatted to carry, for example, information that can be stored or transmitted. The information may include, for example, instructions for performing a method or data generated by one of the described implementations. For example, the signal may be formatted to carry a bit stream of the described example. Such a signal may be formatted as, for example, an electromagnetic wave (e.g., using the radio frequency portion of a spectrum) or a baseband signal. Formatting may include, for example, encoding the data stream and modulating a carrier wave using the encoded data stream. The information carried by the signal may be, for example, analog or digital information. The signal may be transmitted over a variety of different wired or wireless links, as is well known. The signal may be stored on or accessed from a processor-readable medium.

[0135] This document describes numerous examples. Features of the examples may be provided individually or in any combination across various claim classes and types. Further, examples may include one or more of the features, devices, or aspects described individually or in any combination across various claim classes and types. For example, the features described herein may be implemented in a bitstream or signal that includes information generated as described herein. This information may allow a decoder to decode the bitstream, encoder, bitstream, and / or decoder according to any of the described embodiments. For example, the features described herein may be implemented by creating and / or transmitting and / or receiving and / or decoding the bitstream or signal. For example, the features described herein may be implemented using methods, processes, apparatus, media storing instructions, media storing data, or signals. For example, the features described herein may be implemented by a TV, set-top box, cellular phone, tablet, or other electronic device performing decoding. The TV, set-top box, cellular phone, tablet, or other electronic device may display (e.g., using a monitor, screen, or other type of display) the obtained image (e.g., a signal reconstructed from the residual of a video bitstream). The TV, set-top box, cellular phone, tablet, or other electronic device may receive a signal including an encoded image and perform decoding.

[0136] This article describes one or more features associated with decoder-side intra-frame mode derivation (DIMD).

[0137] When DIMD is applied, one or more (e.g., up to five) intra-frame patterns can be derived from reconstructed neighbor samples (e.g., by analyzing the directionality of the content surrounding the current block). Five predictors can be combined with planar pattern predictors (e.g., with weights derived from the histogram of gradients (HoG)). The HoG can be computed on an L-shaped template formed by the reconstructed samples (e.g., an L-shaped template with three sample widths / heights), as... Figure 5 The image is shown in the figure. The template can be created using a Sobel filter (e.g., for...). Figure 5 The DIMD prediction is obtained by accumulating the magnitudes of all gradients at a given direction using samples within the gray area. The direction with the highest accumulated magnitude can be selected as the primary and secondary DIMD modes. Predictors obtained using DIMD modes can be mixed to form a final DIMD prediction. Uniform or spatial mixing can be used (e.g., where DIMD predictors are combined with planar predictors, for example, using weights that depend on the relative magnitudes of modes in the gradient histogram).

[0138] The same lookup table (LUT)-based integerization scheme used by CCLM can be used to perform division in weight derivation. Division in orientation calculation can be described in (4): Orientation = G y / G x (4).

[0139] The division operation in orientation calculation can be performed by the following LUT-based schemes in (5), (6), (7) and (8):

[0140] The following (9) can describe the (e.g., possible) values ​​of DivSigTable: .

[0141] For size W × H The weights of the blocks can be modified for the derived patterns (e.g., each of the five derived patterns), for example, if one of the histogram values ​​above or to the left is twice as large as the other. In this case, the weights can depend on position. The weights can be calculated. The weights can be calculated according to equation (8), for example, if the histogram above is twice the size of the histogram to the left: The weights can be calculated using equation (9), for example, if the left histogram is twice the size of the top histogram: Refer to equations (8) and (9). wDimd i It can be the unmodified uniform weights of DIMD, and Δ i This can be predefined (e.g., set to 10). The weight of the plane is fixed at 21 / 64 (~1 / 3). The remaining weight of 43 / 64 (~2 / 3) can be shared between HoG IPMs (e.g., two HoG IPMs), for example, proportional to the amplitude of their HoG bars, such as... Figure 6 As shown in the diagram.

[0142] Figure 6 An example technique for deriving prediction blocks is illustrated.

[0143] The derived intra-frame modes can be included in a primary list of most probable intra-frame modes (MPMs). The DIMD process can be performed before the MPM list is constructed. The primary derived intra-frame modes of the DIMD block can be stored in blocks. The primary derived intra-frame modes can be used for the construction of MPM lists for adjacent blocks.

[0144] The regions of adjacent reconstructed samples (e.g., for calculating gradient histograms) can depend on the availability of reconstructed samples. The region of the decoded reference sample of the current WxH brightness CB can extend to the upper right (e.g., up to W additional columns if available). The region of the decoded reference sample of the current WxH brightness CB can extend to the lower left (e.g., up to H additional rows if available).

[0145] Figure 7 The illustration depicts an example DIMD process (e.g., sometimes referred to as a regular DIMD process). As illustrated, at 710, adjacent reconstructed reference samples (e.g., as reference regions) can be selected. At 720, HoG can be derived. At 730, intra-frame modes (e.g., up to five intra-frame modes) can be derived. At 740, DIMD fusion weights can be derived. Predictions can then be constructed.

[0146] This paper describes one or more features associated with the Adaptive Reference Region DIMD.

[0147] Figure 8 The diagram illustrates the reference regions for DIMD (e.g., DIMD_TL), DIMD_T, and DIMD_L modes. Providing multiple DIMD modes (e.g., DIMD_T and DIMD_L) allows different neighborhoods to be selected as reference regions. In the DIMD_T mode, the top-left, top, and top-right neighborhoods can be used as reference regions. In the DIMD_L mode, the top-left, left, and bottom-left neighborhoods can be used as reference regions. The number of lines in the reference regions of DIMD_T and DIMD_L modes can be four. The original DIMD mode is called DIMD_TL.

[0148] In the coding unit (CU) (e.g., each CU), signals can be sent to notify things such as (e.g., cu_dimd_mode DIMD mode indications such as (e.g., in DIMD-enabled flags) cu_dimd_flag Then, it is determined which DIMD mode is used. Table 1 summarizes this. cu_dimd_mode Example binarization cu_dimd_mode Binarization name 0 0 DIMD_TL 1 10 DIMD_T 2 11 DIMD_L Table 1: cu_dimd_mode Binarization.

[0149] This article describes one or more features associated with DIMD merging. Figure 9 The illustration shows an example DIMD process based on merged gradient histograms (MHoG).

[0150] If DIMD merging is used, DIMD information extracted from neighboring blocks can be used to compute intra-prediction for the current block. MHoG (e.g., HoG based on neighboring blocks) can be computed for the current block. Neighboring blocks encoded using DIMD or DIMD merging can be considered (e.g., only neighboring blocks encoded using DIMD or DIMD merging) (e.g., as illustrated at 930).

[0151] If a DIMD or DIMD-merged adjacent block (e.g., a single DIMD or DIMD-merged adjacent block) is available, its gradient histogram can be used to form the MHoG for the current block. If more than one DIMD or DIMD-merged adjacent block is available, the corresponding histograms can be combined (e.g., by averaging the bar amplitudes) to derive the MHoG (as illustrated at 940).

[0152] MHoG can be used to compute intra-frame prediction modes and weights (e.g., as in conventional DIMD). Orientation modes and their weights corresponding to the N (e.g., N=5) highest amplitudes in the MHoG can be selected (e.g., as illustrated at 950). Corresponding predictors can be mixed (e.g., as in conventional DIMD), as illustrated at 960.

[0153] In some examples, the DIMD merge mode may be available (e.g., only) if the current block has at least one neighbor encoded using DIMD or DIMD merge (as illustrated at 920). In this case, the use of DIMD merge may be signaled (e.g., using a CU-level flag encoded in CABAC). A CABAC context may be included to support the encoding of the DIMD merge flag. In terms of signaling, DIMD merge may be considered a sub-mode of DIMD. For example, the DIMD merge flag may be signaled (e.g., only) if DIMDflag = 1 and if there are neighboring CUs encoded using DIMD or DIMD merge.

[0154] DIMD merging patterns can take into account spatial characteristics (as in regular DIMD). The process for constructing the MHoG can prioritize neighboring and current histograms based on their respective correlations. Subsets of DIMD histograms and histogram merging can be stored (e.g., CUs encoded using DIMD patterns are reused in the MHoG process).

[0155] Adaptive reference area DIMD merging can be enabled. The merging process of the histogram can be adapted to the characteristics of the histograms of adjacent CUs. One or more features described herein can reduce the memory footprint caused by storing the histograms of CUs encoded using DIMD.

[0156] This paper provides one or more features associated with adaptive reference region DIMD merging. Figure 10 The illustration shows an example of adaptive reference region DIMD merging. As shown, at 1010, the device can recognize and... cu_dimd_merge_ mode The corresponding region R. At 1020, the device can check whether at least one neighbor in R is encoded using DIMD or DIMD merging. If at least one neighbor in R is encoded using DIMD or DIMD merging, then at 1030, the device can select the HoG of the adjacent reconstructed CU encoded in DIMD or DIMD merging within region R. At 1040, the device can derive the MHoG from the selected HoG. At 1050, the device can derive one or more (e.g., up to five) intra-frame modes. At 1060, the device can derive the DIMD fusion weights and construct a prediction.

[0157] The adaptive reference region can be extended to DIMD merging. This can be done for the following instructions (e.g., cu_dimd_merge_ mode Encoding: This instruction indicates whether the MHoG procedure uses a histogram of the CUs encoded in DIMD or DIMD merge located in the left, top-left, or top region. Left and top regions may be considered (e.g., only left and top regions). The left and top regions may or may not include the top-left region. A considered CU encoded in DIMD or DIMD merge mode may not be adjacent to the current CU. For example, a considered CU may be located in a (pre)defined region of size W x H, such as... Figure 11 As shown in the diagram. Figure 11 The illustration shows adjacent blocks encoded in DIMD or DIMD merge in the left, top-left, or top regions.

[0158] The reference region used to calculate the histogram of the current CU can be a regular reference region (e.g., including the left and top samples, e.g., such as...). Figure 5(as illustrated in the diagram). The left sample (e.g., only the left sample) can be used for the reference region (e.g., if...). cu_ dimd_merge_mode (Indicates left). The sample above (e.g., only the sample above) can be used for the reference region (e.g., if...). cu_ dimd_merge_mode (Indicator top).

[0159] This article provides one or more features associated with adaptive merging of histograms.

[0160] MHoG can be implemented by summing the histogram bars of the CUs encoded in DIMD or DIMD merge mode and dividing by the number of merged histograms (e.g., averaging the histograms).

[0161] Figure 12 The diagram illustrates neighboring blocks encoded in DIMD or DIMD merging within the left, top-left, or top regions. Histograms can be implicitly weighted based on the number of samples in neighboring CUs. Histograms can also be weighted based on the proximity of a reference sample to the current CU. It might be desirable to favor histograms calculated using reference samples located close to the current CU. For example, in... Figure 12 In this context, the top reference sample of CU-3 (e.g., R3) can be relatively far from the current CU, while the reference sample of CU-2 (R2) can be relatively close to the current CU.

[0162] The bars of the HoG in R3 can be rescaled using weights. For example, the bars of the HoG in R3 could be rescaled with weights lower than those of the bars of the HoG in R2 (e.g., because the distance D3(top) of the top reference sample in R3 is farther than the distance D2(top) of the top reference sample in R2 (D3>D2)). The distances between the current block and samples in the reference region can be tabulated. The distances can vary (e.g., every four samples). The distance can be the average distance between the top and left reference samples. The distance D3 can be the sum of the distances between the top and left reference samples (D3(top) + D3(left)). The weights can be a function of the distances. For example, the weights can decrease as the value of the distance decreases.

[0163] The MHoG (e.g., which may be derived by averaging adjacent HoGs in the available space (or a location-weighted average, as described herein)) can be mixed with the HoG of the current CU (e.g., via merging). For example, the MHoG and the current HoG can be merged into a final mixed HoG (e.g., using fixed weights, such as weight_MHoG = 1 / 4 and weight_HoG = 3 / 4). Regarding signaling, the DIMD merge flag can be notified without signaling. cu_dimd_merge (For example, because only the final HoG histogram is used).

[0164] This paper provides one or more features associated with the fusion weights of DIMD. The fusion weights (e.g., as shown in the figure) can be derived as a function of the bar amplitudes of the DIMD modes to be blended (e.g., fused). Figure 10 (Illustrated at position 1060 in the diagram). For example, the derived intra-frame mode with the highest amplitude (e.g., two intra-frame modes) can be blended with a weight of 3. Other derived intra-frame modes with the highest amplitude (e.g., three other intra-frame modes) can be blended with a weight of 2. The blending weight can be a function of the bar amplitude and the distance from the current CU.

[0165] This paper provides one or more features associated with DIMD merging without storing HoGs. DIMD merging can utilize significant memory bandwidth (e.g., because the HoG used by CUs encoded in DIMD modes is stored), for example, if many CUs are encoded using a DIMD merging mode. One or more (e.g., up to N, e.g., N=5) intra-modes that are merged using DIMD or a DIMD merging mode can be stored. Storing a subset of intra-modes can reduce the amount of data to be stored per CU. The N (e.g., N=5) intra-modes with the highest occurrence rates can be selected as modes to be merged. The merging weights can be calculated as described herein (e.g., as used in a regular DIMD process or using merging weights based on bar amplitude) or as a function of the occurrence rates of the N intra-modes.

[0166] The fusion weights of intra-frame modes (e.g., each intra-frame mode) can be calculated in proportion to the size of the DIMD block encoded using that intra-frame mode.

[0167] Although features and elements have been described above in specific combinations, those skilled in the art will appreciate that each feature or element may be used individually or in any combination with other features and elements. Furthermore, the methods described herein can be implemented in a computer program, software, or firmware incorporated in a computer-readable medium for execution by a computer or processor. Examples of computer-readable media include electronic signals (transmitted via wired or wireless connections) and computer-readable storage media. Examples of computer-readable storage media include, but are not limited to, read-only memory (ROM), random access memory (RAM), registers, cache memory, semiconductor memory devices, magnetic media (such as internal hard disks and removable disks), magnetic-optical media, and optical media (such as CD-ROMs and digital multifunction discs (DVDs)). The processor associated with the software can be used to implement a radio frequency transceiver for a WTRU, UE, terminal, base station, RNC, or any host computer.

Claims

1. A method for video decoding, the method comprising: Determine whether to enable decoder-side intra-frame mode derivation (DIMD) merging mode for the current block; Determine reference samples to be used for the DIMD merging mode, wherein the reference samples include a first reference sample associated with a first gradient histogram and a second reference sample associated with a second gradient histogram; The first gradient histogram is expanded or reduced based on the first distance between the first reference sample and the current block; The second gradient histogram is expanded or reduced based on the second distance between the second reference sample and the current block; The merged gradient histogram is derived based on the expanded and contracted first gradient histogram and the expanded and contracted second gradient histogram. One or more intra-frame prediction modes are derived based on merged gradient histograms. The prediction of the current block is obtained based on the one or more intra-frame prediction modes; as well as The current block is decoded based on the prediction.

2. The method of claim 1, wherein scaling the first gradient histogram based on the first distance between the first reference sample and the current block comprises: The first scaling factor is determined based on the first distance between the first reference sample and the current block; as well as The first gradient histogram is expanded or reduced using the first scaling factor. Furthermore, the expansion and contraction of the second gradient histogram based on the second distance between the second reference sample and the current block includes: The second scaling factor is determined based on the second distance between the second reference sample and the current block; as well as The second gradient histogram is expanded or reduced using the second expansion / reduction factor.

3. The method of claim 2, wherein the first scaling factor is inversely proportional to the first distance, and the second scaling factor is inversely proportional to the second distance.

4. The method of any one of claims 1-3, wherein the first distance between the first reference sample and the current block is: The sum of the vertical distance between the first reference sample and the current block and the horizontal distance between the first reference sample and the current block; or The average of the vertical distance between the first reference sample and the current block and the horizontal distance between the first reference sample and the current block.

5. The method of any one of claims 1-4, wherein obtaining the prediction of the current block based on the one or more intra-frame prediction modes comprises: Determine the first fusion weight for the first intra-frame prediction mode associated with the highest frequency bar of the merged histogram; Determine the second fusion weight for the second intra-frame prediction corresponding to the second N2 highest bar of the merged histogram; as well as The prediction of the current block is derived by mixing the first intra-frame prediction mode with the first fusion weight and mixing the second intra-frame prediction mode with the second fusion weight.

6. The method of any one of claims 1-5, wherein deriving one or more intra-frame prediction modes based on the merged gradient histogram comprises: Determine the first intra-frame prediction mode associated with the highest frequency bar in the merged gradient histogram; as well as Determine the second intra-frame prediction mode associated with the second high-frequency bars in the merged gradient histogram.

7. A method for video encoding, the method comprising: Determine whether to enable decoder-side intra-frame mode derivation (DIMD) merging mode for the current block; Determine reference samples to be used for the DIMD merging mode, wherein the reference samples include a first reference sample associated with a first gradient histogram and a second reference sample associated with a second gradient histogram; The first gradient histogram is expanded or reduced based on the first distance between the first reference sample and the current block; The second gradient histogram is expanded or reduced based on the second distance between the second reference sample and the current block; The merged gradient histogram is derived based on the expanded and contracted first gradient histogram and the expanded and contracted second gradient histogram. as well as The current block is encoded based on the merged gradient histogram.

8. The method of claim 7, wherein scaling the first gradient histogram based on the first distance between the first reference sample and the current block comprises: The first scaling factor is determined based on the first distance between the first reference sample and the current block; as well as The first gradient histogram is expanded or reduced using the first scaling factor. Furthermore, the expansion and contraction of the second gradient histogram based on the second distance between the second reference sample and the current block includes: The second scaling factor is determined based on the second distance between the second reference sample and the current block; as well as The second gradient histogram is expanded or reduced using the second expansion / reduction factor.

9. The method of claim 8, wherein the first scaling factor is inversely proportional to the first distance, and the second scaling factor is inversely proportional to the second distance.

10. The method of any one of claims 7-9, wherein the reference sample further comprises a third reference sample associated with the first gradient histogram, and scaling the first gradient histogram based on a first distance between the first reference sample and the current block comprises: Determine the third distance between the third reference sample and the current block; The first distance and the third distance are added together or averaged to determine the fourth distance; as well as The first gradient histogram is expanded or reduced based on the fourth distance.

11. The method of any one of claims 7-10, wherein deriving the merged gradient histogram based on the expanded first gradient histogram and the expanded second gradient histogram comprises: The first fusion weight is determined based on the first bar amplitude associated with the expanded first gradient histogram, and the second fusion weight is determined based on the second bar amplitude associated with the expanded second gradient histogram. as well as The expanded first gradient histogram and the expanded second gradient histogram are blended based on the first fusion weight and the second fusion weight.

12. The method of any one of claims 7-11, wherein encoding the current block based on the merged gradient histogram comprises: One or more intra-frame prediction modes are derived based on merged gradient histograms. The prediction of the current block is obtained based on the one or more intra-frame prediction modes; as well as The current block is encoded based on the prediction.

13. A video decoding device, the device comprising: The processor is configured as follows: Determine whether to enable decoder-side intra-frame mode derivation (DIMD) merging mode for the current block; Determine reference samples to be used for the DIMD merging mode, wherein the reference samples include a first reference sample associated with a first gradient histogram and a second reference sample associated with a second gradient histogram; The first gradient histogram is expanded or reduced based on the first distance between the first reference sample and the current block; The second gradient histogram is expanded or reduced based on the second distance between the second reference sample and the current block; The merged gradient histogram is derived based on the expanded and contracted first gradient histogram and the expanded and contracted second gradient histogram. as well as The current block is decoded based on the merged gradient histogram.

14. The apparatus of claim 13, wherein the processor is configured to expand or shrink the first gradient histogram based on a first distance between the first reference sample and the current block, the processor is configured to: The first scaling factor is determined based on the first distance between the first reference sample and the current block; as well as The first gradient histogram is expanded or reduced using the first scaling factor, and The processor is configured to expand or shrink the second gradient histogram based on a second distance between the second reference sample and the current block, including that the processor is configured to: The second scaling factor is determined based on the second distance between the second reference sample and the current block; as well as The second gradient histogram is expanded or reduced using the second expansion / reduction factor.

15. The device of claim 14, wherein the first scaling factor is inversely proportional to the first distance, and the second scaling factor is inversely proportional to the second distance.

16. The device of any one of claims 13-15, wherein the reference sample further comprises a third reference sample associated with the first gradient histogram, and the processor is configured to scale the first gradient histogram based on a first distance between the first reference sample and the current block, wherein the processor is configured to: Determine the third distance between the third reference sample and the current block; The first distance and the third distance are added together or averaged to determine the fourth distance; as well as The first gradient histogram is expanded or reduced based on the fourth distance.

17. The device of any one of claims 13-16, wherein deriving the merged gradient histogram based on the processor configured to do so based on the expanded first gradient histogram and the expanded second gradient histogram includes the processor being configured to: A first fusion weight is determined based on the amplitude of a first bar associated with a first scaled gradient histogram, and a second fusion weight is determined based on the amplitude of a second bar associated with a second scaled gradient histogram; and The expanded first gradient histogram and the expanded second gradient histogram are blended based on the first fusion weight and the second fusion weight.

18. The device of any one of claims 13-17, wherein the processor is configured to decode the current block based on a merged gradient histogram, the processor is configured to: One or more intra-frame prediction modes are derived based on merged gradient histograms. The prediction of the current block is obtained based on the one or more intra-frame prediction modes; as well as The current block is decoded based on the prediction.

19. A video encoding device, comprising: The processor is configured as follows: Determine whether to enable decoder-side intra-frame mode derivation (DIMD) merging mode for the current block; Determine reference samples to be used for the DIMD merging mode, wherein the reference samples include a first reference sample associated with a first gradient histogram and a second reference sample associated with a second gradient histogram; The first gradient histogram is expanded or reduced based on the first distance between the first reference sample and the current block; The second gradient histogram is expanded or reduced based on the second distance between the second reference sample and the current block; The merged gradient histogram is derived based on the expanded and contracted first gradient histogram and the expanded and contracted second gradient histogram. as well as The current block is encoded based on the merged gradient histogram.

20. The apparatus of claim 19, wherein the processor is configured to expand or shrink the first gradient histogram based on a first distance between the first reference sample and the current block, the processor is configured to: The first scaling factor is determined based on the first distance between the first reference sample and the current block; as well as The first gradient histogram is expanded or reduced using the first scaling factor, and The processor is configured to expand or shrink the second gradient histogram based on a second distance between the second reference sample and the current block, including that the processor is configured to: The second scaling factor is determined based on the second distance between the second reference sample and the current block; as well as The second gradient histogram is expanded or reduced using the second expansion / reduction factor.