Superpixel spatial merging for implicit neural representations
By using implicit neural representation (INR) for superpixel spatial merging and optimizing inter-segment distance and coding performance, the problem of low coding efficiency in video coding systems is solved, achieving more efficient video signal compression and transmission.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-21
- Publication Date
- 2026-04-07
AI Technical Summary
Existing video coding systems struggle to effectively utilize the distance between video partitions and optimize coding performance when compressing and transmitting digital video signals, resulting in low coding efficiency.
Implicit Neural Representation (INR) is used for superpixel spatial merging. Techniques such as greedy search, brute force, reinforcement learning, or genetic algorithms are used to determine partition groups. The trained model is then used to generate partition group predictions to optimize encoding performance.
It improves the efficiency and quality of video encoding, and reduces storage and transmission bandwidth requirements through reasonable partitioning and the application of encoding technology.
Smart Images

Figure CN121816740A_ABST
Abstract
Description
Cross-references to related applications
[0001] This application claims the benefit of European Patent Application No. 23306407.0, filed on 24 August 2023, the entire contents of which are incorporated herein by reference. Background Technology
[0002] Video coding systems can be used to compress digital video signals, for example, to reduce the storage and / or transmission bandwidth required for such signals. Video coding systems can include, for example, block-based, wavelet-based, and / or object-based systems. Summary of the Invention
[0003] Systems, methods, and instrumentality for performing superpixel spatial merging for implicit neural representations (INRs) are disclosed. Signals can be encoded, for example, using partitions and INRs. Merging or grouping of portions of the signal can be supported and / or permitted. Information from the merging of these portions can be used to train representation parameters (e.g., INR parameters) for the grouped portions of the signal. This information can be encoded in video data (e.g., a bitstream). The video data (e.g., encoded video data) can be decoded to determine this information.
[0004] Devices (e.g., video encoders or video decoders) can be configured to perform the partitioning and merging of signals. For example, the device can determine partitions (e.g., a first partition and / or a second partition) associated with an input signal. The device can determine partition groups. The device can determine a first partition group and / or a second partition group. For example, a first partition group may include a first partition. A second partition group may include a second partition. Partitions to be included in the partition group can be determined. For example, partitions to be included in the partition group can be determined based on the distance between partitions (e.g., the distance between the first and second partitions). Partitions to be included in the partition group can be determined based on the effects associated with partition grouping. The effects associated with partition grouping can be determined, for example, based on optimization of the coding performance associated with partition grouping. Optimization of the coding performance associated with partition grouping can be associated with using one or more of the following: greedy search, brute force methods, reinforcement learning, genetic algorithms, etc. A model can be used to generate partition grouping predictions. The model can be trained. The device can determine representation parameters associated with the partition group (e.g., determining a first representation parameter associated with the first partition group and / or determining a second representation parameter associated with the second partition group). The device can determine the encoding technique associated with a partition group (e.g., a first encoding technique associated with a first partition group and / or a second encoding technique associated with a second partition group). The device can encode information associated with the partition, such as one or more of the following: first partition, second partition, first partition group, second partition group, first representation parameter, second representation parameter, first encoding technique associated with the first partition group, second encoding technique associated with the second partition group, etc. The device can include encoded information (e.g., encoded first partition, second partition, first partition group, second partition group, first representation parameter, second representation parameter, first encoding technique associated with the first partition group, second encoding technique associated with the second partition group, etc.) in the video data (e.g., bitstream). The device can include an indication of using an encoding group unit in the video data (e.g., bitstream). The device can include a group information index in the video data (e.g., bitstream). The group information index can indicate group information associated with a partition (e.g., first partition and / or second partition).
[0005] Devices (e.g., video encoders or video decoders) can be configured to perform the partitioning and merging of signals. The device can acquire an input signal. The device can determine partitions for the input signal. The device can determine groups for the determined partitions. The device can determine the corresponding representation parameters for each partition group. The device can determine the encoding technique associated with the partition group. The device can encode the partitions, partition groups, and the representation parameters associated with each partition group. The device can include the encoded partitions, partition groups, and the representation parameters associated with each partition group in the video data. The video data can include the determined encoding technique associated with the partition group. The video data can include instructions indicating the use of coding unit groups.
[0006] Partition groups can be determined. For example, partition groups can be determined based on the distance between partitions. Partition groups can be determined based on the impact of grouping on coding performance (e.g., based on using greedy search, brute force methods, reinforcement learning, genetic algorithms, etc.). Partition groups can be determined based on a trained model (e.g., a machine learning model). The device can train the model to generate partition grouping predictions.
[0007] Devices (e.g., video encoding or decoding devices) can decode video data to determine partitioning information. Partitioning information may indicate partitions, partition groups, and corresponding group coding technology information associated with each partition group. The device can determine the decoding technology for the partition group based on the partitioning information. The device can determine the signal values for the partition group based on the determined decoding technology. The device can reconstruct the signal, for example, based on the signal values.
[0008] The systems, methods, and tools described herein may relate to decoders. In some examples, the systems, methods, and tools described herein may relate to encoders. In some examples, the systems, methods, and tools described herein may relate to signals (e.g., from an encoder and / or received by a decoder). Computer-readable media may include instructions for causing one or more processors to perform the methods described herein. A computer program product may include instructions that, when executed by one or more processors, cause one or more processors to perform the methods described herein. Attached Figure Description
[0009] Figure 1A This is a system diagram illustrating an exemplary communication system that can implement one or more of the disclosed embodiments.
[0010] Figure 1B The illustration shows that, according to the embodiment, it can be used... Figure 1A The diagram shows a system diagram of an exemplary wireless transmit / receive unit (WTRU) used within a communication system.
[0011] Figure 1C The illustration shows that, according to the embodiment, it can be used... Figure 1A The diagram shows an exemplary radio access network (RAN) and an exemplary core network (CN) used within a communication system.
[0012] Figure 1D The illustration shows that, according to the embodiment, it can be used... Figure 1A The diagram shows another exemplary RAN and another exemplary CN used within the communication system.
[0013] Figure 2 An exemplary video encoder is illustrated.
[0014] Figure 3 An exemplary video decoder is illustrated.
[0015] Figure 4 An exemplary system that can be implemented in various aspects and examples is illustrated.
[0016] Figure 5 An exemplary simple neural network for implicit neural representation (INR) is illustrated.
[0017] Figure 6 An exemplary process for encoding signals using INR is illustrated.
[0018] Figure 7 An exemplary encoding process is illustrated.
[0019] Figure 8 An exemplary encoding process for sequential encoding input is illustrated.
[0020] Figure 9 An exemplary process for generating the reconstructed signal is illustrated.
[0021] Figure 10 An exemplary result implemented on the image is illustrated.
[0022] Figure 11 The illustration shows an example of superpixel segmentation before merging.
[0023] Figure 12 The illustration shows an example of merged superpixel segmentation. Detailed Implementation
[0024] A more detailed understanding can be obtained from the following exemplary description taken in conjunction with the accompanying drawings.
[0025] Figure 1AThis diagram illustrates an exemplary communication system 100 that may implement one or more of the disclosed embodiments. The communication system 100 may be a multiple access system that provides content (such as voice, data, video, messaging, broadcasting, etc.) to multiple wireless users. The communication system 100 enables multiple wireless users to access such content by sharing system resources, including wireless bandwidth. For example, the communication system 100 may employ one or more channel access methods, such as Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Frequency Division Multiple Access (FDMA), Orthogonal FDMA (OFDMA), Single Carrier FDMA (SC-FDMA), Zero-Tail Unique Word DFT Extended OFDM (ZT UW DTS-s OFDM), Unique Word OFDM (UW-OFDM), Resource Block Filtered OFDM, Filter Bank Multicarrier (FBMC), etc.
[0026] like Figure 1A As shown, the communication system 100 may include wireless transmit / receive units (WTRUs) 102a, 102b, 102c, 102d, RAN 104 / 113, CN 106 / 115, public switched telephone network (PSTN) 108, Internet 110, and other networks 112. However, it should be understood that the disclosed embodiments contemplate any number of WTRUs, base stations, networks, and / or network elements. Each of the WTRUs 102a, 102b, 102c, 102d may be any type of device configured to operate and / or communicate in a wireless environment. As an example, WTRUs 102a, 102b, 102c, and 102d (any of which may be referred to as a “station” and / or “STA”) may be configured to transmit and / or receive wireless signals and may include user equipment (UE), mobile stations, fixed or mobile subscriber units, subscription-based units, pagers, cellular phones, personal digital assistants (PDAs), smartphones, laptops, netbooks, personal computers, wireless sensors, hotspots or Mi-Fi devices, Internet of Things (IoT) devices, watches or other wearable devices, head-mounted displays (HMDs), vehicles, drones, medical devices and applications (e.g., remote surgery), industrial devices and applications (e.g., robots and / or other wireless devices operating in industrial and / or automated processing chain environments), consumer electronics devices, devices operating on commercial and / or industrial wireless networks, etc. Any of WTRUs 102a, 102b, 102c, and 102d may be interchangeably referred to as a UE.
[0027] The communication system 100 may also include base station 114a and / or base station 114b. Each of base stations 114a and 114b may be any type of device configured to wirelessly interface with at least one of WTRUs 102a, 102b, 102c, and 102d to facilitate access to one or more communication networks (such as CN 106 / 115, Internet 110, and / or other networks 112). As an example, base stations 114a and 114b may be base transceiver stations (BTS), Node-B, eNode B, home Node B, home eNode B, gNB, NR Node B, site controllers, access points (APs), wireless routers, etc. Although base stations 114a and 114b are each depicted as a single element, it should be understood that base stations 114a and 114b may include any number of interconnected base stations and / or network elements.
[0028] Base station 114a may be part of RAN 104 / 113, which may also include other base stations and / or network elements (not shown), such as base station controllers (BSCs), radio network controllers (RNCs), relay nodes, etc. Base station 114a and / or base station 114b may be configured to transmit and / or receive radio signals on one or more carrier frequencies (which may be referred to as cells (not shown)). These frequencies may be in licensed spectrum, unlicensed spectrum, or a combination of licensed and unlicensed spectrum. Cells may provide coverage for a specific geographic area that may be relatively fixed or change over time. Cells may also be divided into cell sectors. For example, the cell associated with base station 114a may be divided into three sectors. Thus, in one embodiment, base station 114a may include three transceivers, one for each cell sector. In embodiments, base station 114a may employ multiple-input multiple-output (MIMO) technology and may utilize multiple transceivers in each cell sector. For example, beamforming may be used to transmit and / or receive signals in a desired spatial direction.
[0029] Base stations 114a and 114b can communicate with one or more of WTRUs 102a, 102b, 102c, and 102d via air interface 116. Air interface 116 can be any suitable wireless communication link (e.g., radio frequency (RF), microwave, centimeter wave, micrometer wave, infrared (IR), ultraviolet (UV), visible light, etc.). Any suitable radio access technology (RAT) can be used to establish air interface 116.
[0030] More specifically, as described above, the communication system 100 can be a multiple access system and can employ one or more channel access schemes, such as CDMA, TDMA, FDMA, OFDMA, SC-FDMA, etc. For example, base stations 114a and WTRUs 102a, 102b, and 102c in RAN 104 / 113 can implement radio technologies such as Universal Mobile Telecommunications System (UMTS) Terrestrial Radio Access (UTRA), and UTRA can use Wideband CDMA (WCDMA) to establish air interfaces 115 / 116 / 117. WCDMA can include communication protocols such as High-Speed Packet Access (HSPA) and / or Evolved HSPA (HSPA+). HSPA can include High-Speed Downlink (DL) Packet Access (HSDPA) and / or High-Speed UL Packet Access (HSUPA).
[0031] In the embodiment, base station 114a and WTRUs 102a, 102b, 102c can implement radio technologies such as Evolved UMTS Terrestrial Radio Access (E-UTRA), which can use Long Term Evolution (LTE) and / or LTE-Advanced (LTE-A) and / or LTE-Advanced Pro (LTE-A Pro) to establish air interface 116.
[0032] In the embodiment, base station 114a and WTRUs 102a, 102b, 102c may implement radio technologies such as NR radio access, which can use a new radio (NR) to establish an air interface 116.
[0033] In the embodiments, base station 114a and WTRUs 102a, 102b, and 102c can implement various radio access technologies. For example, base station 114a and WTRUs 102a, 102b, and 102c can implement both LTE radio access and NR radio access, for example, using the dual connectivity (DC) principle. Therefore, the air interface utilized by WTRUs 102a, 102b, and 102c can be characterized by various types of radio access technologies and / or transmissions to / from various types of base stations (e.g., eNBs and gNBs).
[0034] In other embodiments, base station 114a and WTRUs 102a, 102b, 102c may implement radio technologies such as IEEE 802.11 (i.e., Wi-Fi), IEEE 802.16 (i.e., WiMAX), CDMA2000, CDMA2000 1X, CDMA2000 EV-DO, Provisional Standard 2000 (IS-2000), Provisional Standard 95 (IS-95), Provisional Standard 856 (IS-856), Global System for Mobile Communications (GSM), Enhanced Data Rate GSM Evolution (EDGE), and GSM EDGE (GERAN).
[0035] Figure 1A Base station 114b can be, for example, a wireless router, home Node B, home eNode B, or access point, and can utilize any suitable RAT to facilitate wireless connectivity in local areas such as business premises, homes, vehicles, campuses, industrial facilities, air corridors (e.g., for drones), roads, etc. In one embodiment, base station 114b and WTRUs 102c, 102d can implement radio technologies such as IEEE 802.11 to establish a wireless local area network (WLAN). In another embodiment, base station 114b and WTRUs 102c, 102d can implement radio technologies such as IEEE 802.15 to establish a wireless personal area network (WPAN). In yet another embodiment, base station 114b and WTRUs 102c, 102d can utilize cellular-based RATs (e.g., WCDMA, CDMA2000, GSM, LTE, LTE-A, LTE-A Pro, NR, etc.) to establish picocells or femtocells. Figure 1A As shown, base station 114b can have a direct connection to Internet 110. Therefore, base station 114b does not need to access Internet 110 via CN 106 / 115.
[0036] RAN 104 / 113 can communicate with CN 106 / 115, which can be any type of network configured to provide voice, data, application, and / or Voice over Internet Protocol (VoIP) services to one or more WTRUs 102a, 102b, 102c, and 102d. Data can have different Quality of Service (QoS) requirements, such as different throughput requirements, latency requirements, fault tolerance requirements, reliability requirements, data throughput requirements, mobility requirements, etc. CN 106 / 115 can provide call control, billing services, location-based services, prepaid calling, internet connectivity, video distribution, etc., and / or perform advanced security functions such as user authentication. Although in Figure 1AAlthough not shown, it should be understood that RAN 104 / 113 and / or CN 106 / 115 can communicate directly or indirectly with other RANs that use the same RAT as or a different RAT than RAN 104 / 113. For example, in addition to connecting to RAN 104 / 113, which can use NR radio technology, CN 106 / 115 can also communicate with another RAN (not shown) that uses GSM, UMTS, CDMA 2000, WiMAX, E-UTRA, or WiFi radio technology.
[0037] CN 106 / 115 can also act as a gateway for WTRU 102a, 102b, 102c, 102d to access PSTN 108, Internet 110, and / or other networks 112. PSTN 108 may include a circuit-switched telephone network providing Common Old-Style Telephone Service (POTS). Internet 110 may include a global system of interconnected computer networks and devices using common communication protocols such as Transmission Control Protocol (TCP), User Datagram Protocol (UDP), and / or Internet Protocol (IP) (in the TCP / IP Internet Protocol suite). Network 112 may include wired and / or wireless communication networks owned and / or operated by other service providers. For example, network 112 may include another CN connected to one or more RANs using the same RAT as or a different RAT than RAN 104 / 113.
[0038] Some or all of the WTRUs 102a, 102b, 102c, and 102d in the communication system 100 may include multi-mode capability (e.g., WTRUs 102a, 102b, 102c, and 102d may include multiple transceivers for communicating with different wireless networks via different wireless links). For example, Figure 1A The WTRU 102c shown can be configured to communicate with a base station 114a that can employ cellular-based radio technology and with a base station 114b that can employ IEEE 802 radio technology.
[0039] Figure 1B This is a system diagram illustrating an exemplary WTRU 102. (Example:) Figure 1B As shown, WTRU 102 may include a processor 118, a transceiver 120, a transmit / receive element 122, a speaker / microphone 124, a keyboard 126, a display / touchpad 128, non-removable memory 130, removable memory 132, a power supply 134, a global positioning system (GPS) chipset 136, and / or other peripheral devices 138, etc. It should be understood that WTRU 102 may include any sub-combination of the foregoing elements while remaining consistent with the embodiments.
[0040] Processor 118 can be a general-purpose processor, a special-purpose processor, a conventional processor, a digital signal processor (DSP), multiple microprocessors, one or more microprocessors associated with a DSP core, a controller, a microcontroller, an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) circuit, any other type of integrated circuit (IC), a state machine, etc. Processor 118 can perform signal encoding, data processing, power control, input / output processing, and / or any other functions that enable WTRU 102 to operate in a wireless environment. Processor 118 can be coupled to transceiver 120, and transceiver 120 can be coupled to transmitting / receiving element 122. Although Figure 1B The processor 118 and transceiver 120 are depicted as separate components, but it should be understood that the processor 118 and transceiver 120 may be integrated together in an electronic package or chip.
[0041] Transmitting / receiving element 122 can be configured to transmit signals to or receive signals from a base station (e.g., base station 114a) via air interface 116. For example, in one embodiment, transmitting / receiving element 122 can be an antenna configured to transmit and / or receive RF signals. In another embodiment, transmitting / receiving element 122 can be a transmitter / detector configured to, for example, transmit and / or receive IR, UV, or visible light signals. In yet another embodiment, transmitting / receiving element 122 can be configured to transmit and / or receive both RF and optical signals. It should be understood that transmitting / receiving element 122 can be configured to transmit and / or receive any combination of wireless signals.
[0042] Although the transmitting / receiving element 122 is in Figure 1B While depicted as a single element, WTRU 102 may include any number of transmit / receive elements 122. More specifically, WTRU 102 may employ MIMO technology. Thus, in one embodiment, WTRU 102 may include two or more transmit / receive elements 122 (e.g., multiple antennas) for transmitting and receiving wireless signals via air interface 116.
[0043] Transceiver 120 can be configured to modulate signals transmitted by transmitting / receiving element 122 and demodulate signals received by transmitting / receiving element 122. As described above, WTRU 102 can have multimode capability. Therefore, transceiver 120 can include multiple transceivers for enabling WTRU 102 to communicate via various RATs (e.g., NR and IEEE 802.11).
[0044] The processor 118 of WTRU 102 can be coupled to a speaker / microphone 124, a keyboard 126, and / or a display / touchpad 128 (e.g., a liquid crystal display (LCD) unit or an organic light-emitting diode (OLED) display unit) and can receive user input data therefrom. The processor 118 can also output user data to the speaker / microphone 124, keyboard 126, and / or display / touchpad 128. Furthermore, the processor 118 can access and store information from any type of suitable memory, such as non-removable memory 130 and / or removable memory 132. Non-removable memory 130 may include random access memory (RAM), read-only memory (ROM), a hard disk, or any other type of memory storage device. Removable memory 132 may include a user identity module (SIM) card, memory stick, secure digital storage (SD) card, etc. In other embodiments, the processor 118 can access and store information from memory that is not physically located on WTRU 102 (such as on a server or home computer (not shown)).
[0045] The processor 118 can receive power from the power supply 134 and can be configured to distribute power to and / or control that power to other components in the WTRU 102. The power supply 134 can be any suitable device for powering the WTRU 102. For example, the power supply 134 may include one or more dry cell batteries (e.g., nickel-cadmium (NiCd), nickel-zinc (NiZn), nickel metal hydride (NiMH), lithium-ion (Li-ion), etc.), solar cells, fuel cells, etc.
[0046] The processor 118 may also be coupled to a GPS chipset 136, which may be configured to provide location information (e.g., longitude and latitude) about the current location of the WTRU 102. In addition to or instead of the information from the GPS chipset 136, the WTRU 102 may receive location information from base stations (e.g., base stations 114a, 114b) via the air interface 116 and / or determine its location based on signal timing received from two or more nearby base stations. It should be understood that the WTRU 102 may acquire location information using any suitable location determination method while remaining consistent with the embodiments.
[0047] The processor 118 may be further coupled to other peripheral devices 138, which may include one or more software and / or hardware modules providing additional features, functions, and / or wired or wireless connectivity. For example, peripheral devices 138 may include accelerometers, electronic compasses, satellite transceivers, digital cameras (for photos and / or video), Universal Serial Bus (USB) ports, vibration devices, television transceivers, hands-free headsets, Bluetooth® modules, FM radio units, digital music players, media players, video game player modules, internet browsers, virtual reality and / or augmented reality (VR / AR) devices, activity trackers, etc. Peripheral devices 138 may include one or more sensors, which may be one or more gyroscopes, accelerometers, Hall effect sensors, magnetometers, orientation sensors, proximity sensors, temperature sensors, time sensors; geolocation sensors; altimeters, light sensors, touch sensors, magnetometers, barometers, gesture sensors, biometric sensors, and / or humidity sensors.
[0048] WTRU 102 may include a full-duplex radio for which the transmission and reception of some or all signals (e.g., associated with specific subframes for both UL (e.g., for transmission) and downlink (e.g., for reception)) may be concurrent and / or simultaneous. The full-duplex radio may include an interference management unit that reduces or substantially eliminates self-interference through hardware (e.g., chokes) or through signal processing by a processor (e.g., a separate processor (not shown) or via processor 118). In embodiments, WTRU 102 may include a half-duplex radio for the transmission and reception of some or all signals (e.g., associated with specific subframes for UL (e.g., for transmission) or downlink (e.g., for reception)).
[0049] Figure 1C This is a system diagram illustrating RAN 104 and CN 106 according to an embodiment. As described above, RAN 104 can communicate with WTRUs 102a, 102b, and 102c via air interface 116 using E-UTRA radio technology. RAN 104 can also communicate with CN 106.
[0050] RAN 104 may include eNode-Bs 160a, 160b, and 160c; however, it should be understood that RAN 104 may include any number of eNode-Bs while remaining consistent with the embodiments. eNode-Bs 160a, 160b, and 160c may each include one or more transceivers for communicating with WTRUs 102a, 102b, and 102c via air interface 116. In one embodiment, eNode-Bs 160a, 160b, and 160c may implement MIMO technology. Therefore, for example, eNode-B 160a may use multiple antennas to transmit and / or receive radio signals from WTRU 102a.
[0051] Each of the eNode-B 160a, 160b, and 160c can be associated with a specific cell (not shown) and can be configured to handle radio resource management decisions, handover decisions, user scheduling in UL and / or DL, etc. Figure 1C As shown, eNode-B 160a, 160b, and 160c can communicate with each other via the X2 interface.
[0052] Figure 1C The CN 106 shown may include a Mobility Management Entity (MME) 162, a Serving Gateway (SGW) 164, and a Packet Data Network (PDN) Gateway (or PGW) 166. Although each of the foregoing elements is depicted as part of CN 106, it should be understood that any of these elements may be owned and / or operated by an entity other than a CN operator.
[0053] The MME 162 can connect to each eNode-B 162a, 162b, 162c in RAN 104 via the S1 interface and can act as a control node. For example, the MME 162 can be responsible for authenticating users of WTRUs 102a, 102b, 102c, bearer activation / deactivation, selecting specific service gateways, etc., during the initial attachment of WTRUs 102a, 102b, 102c. The MME 162 can provide control plane functions for handover between RAN 104 and other RANs (not shown) employing other radio technologies such as GSM and / or WCDMA.
[0054] The SGW 164 can connect to each eNode B 160a, 160b, or 160c in RAN 104 via the S1 interface. The SGW 164 typically routes and forwards user data packets to / from WTRUs 102a, 102b, or 102c. The SGW 164 can perform other functions, such as anchoring the user plane during inter-eNode B handover, triggering paging when DL data is available for WTRUs 102a, 102b, or 102c, and managing and storing the context of WTRUs 102a, 102b, or 102c.
[0055] SGW 164 can connect to PGW 166, which can provide WTRU 102a, 102b, and 102c with access to packet-switched networks (such as Internet 110) to facilitate communication between WTRU 102a, 102b, and 102c and IP-enabled devices.
[0056] CN 106 can facilitate communication with other networks. For example, CN 106 can provide WTRUs 102a, 102b, and 102c with access to circuit-switched networks (such as PSTN 108) to facilitate communication between WTRUs 102a, 102b, and 102c and traditional landline communication equipment. For example, CN 106 may include or be able to communicate with an IP gateway (e.g., an IP Multimedia Subsystem (IMS) server), which acts as an interface between CN 106 and PSTN 108. Furthermore, CN 106 can provide WTRUs 102a, 102b, and 102c with access to other networks 112, which may include other wired and / or wireless networks owned and / or operated by other service providers.
[0057] Despite WTRU in Figures 1A to 1D The term is described as a wireless terminal, but it is conceivable that in some representative embodiments, such a terminal may use (e.g., temporarily or permanently) a wired communication interface with a communication network.
[0058] In a representative embodiment, the other network 112 may be a WLAN.
[0059] In Infrastructure Basic Services Set (BSS) mode, a WLAN may have an access point (AP) for the BSS and one or more stations (STAs) associated with the AP. The AP may have access or an interface to a distribution system (DS) or another type of wired / wireless network that carries traffic flows to and from the BSS. Traffic flows originating outside the BSS and destined for a STA can arrive at and be delivered to the STA via the AP. Traffic flows originating from a STA and destined for a destination outside the BSS can be sent to the AP for delivery to the appropriate destination. Traffic flows between STAs within the BSS can be transmitted via the AP, for example, where a source STA can send a traffic flow to the AP, and the AP can deliver the traffic flow to the destination STA. Traffic flows between STAs within the BSS can be considered and / or referred to as peer-to-peer traffic flows. Peer-to-peer traffic flows can be transmitted between a source STA and a destination STA (e.g., directly) via Direct Link Establishment (DLS). In some representative embodiments, the DLS may use 802.11e DLS or 802.11z Tunneled DLS (TDLS). WLANs using the Standalone BSS (IBSS) mode can function without access points (APs), and STAs within the IBSS or using the IBSS (e.g., all STAs) can communicate directly with each other. The IBSS communication mode may sometimes be referred to as the "ad-hoc" communication mode in this document.
[0060] When operating in 802.11ac infrastructure mode or a similar mode, the AP can transmit beacons on a fixed channel, such as the primary channel. The primary channel can be of fixed width (e.g., a 20 MHz wide bandwidth) or dynamically set via signaling. The primary channel can be the operating channel of the BSS and can be used by the STA to establish a connection with the AP. In some representative embodiments, such as in an 802.11 system, Carrier Sense Multiple Access with Collision Avoidance (CSMA / CA) can be implemented. With CSMA / CA, each STA (including the AP) can listen on the primary channel. If a particular STA listens / detects and / or determines that the primary channel is busy, that particular STA can back off. In a given BSS, at any given time, there can be only one STA (e.g., only one station) transmitting.
[0061] High-throughput (HT) STAs can communicate using a 40 MHz wide channel, for example, by combining a primary 20 MHz channel with adjacent or non-adjacent 20 MHz channels to form a 40 MHz wide channel.
[0062] Very High Throughput (VHT) STAs can support wide channels of 20 MHz, 40 MHz, 80 MHz, and / or 160 MHz. 40 MHz and / or 80 MHz channels can be formed by combining consecutive 20 MHz channels. A 160 MHz channel can be formed by combining eight consecutive 20 MHz channels, or by combining two non-consecutive 80 MHz channels (which can be referred to as an 80+80 configuration). For the 80+80 configuration, the channel-coded data can be divided into two streams by a segmented parser. Inverse Fast Fourier Transform (IFFT) processing and time-domain processing can be performed separately on each stream. These streams can be mapped onto two 80 MHz channels, and the data can be transmitted by the transmitting STA. At the receiver of the receiving STA, the operations described above for the 80+80 configuration can be reversed, and the combined data can be sent to the Media Access Control (MAC).
[0063] 802.11af and 802.11ah support sub-1 GHz operating modes. Compared to those used in 802.11n and 802.11ac, 802.11af and 802.11ah have reduced channel operating bandwidth and carriers. 802.11af supports 5 MHz, 10 MHz, and 20 MHz bandwidths in the TV white space (TVWS) spectrum, while 802.11ah uses non-TVWS spectrum to support 1 MHz, 2 MHz, 4 MHz, 8 MHz, and 16 MHz bandwidths. According to representative embodiments, 802.11ah can support instrument-type control / machine-type communications, such as MTC devices, in macro coverage areas. MTC devices may have certain capabilities, such as limited capabilities, including support (e.g., only support) certain and / or limited bandwidths. MTC devices may include batteries with a battery life exceeding a threshold (e.g., to maintain a very long battery life).
[0064] WLAN systems supporting multiple channels and channel bandwidths (such as 802.11n, 802.11ac, 802.11af, and 802.11ah) include channels that can be designated as primary channels. The primary channel can have a bandwidth equal to the maximum common operating bandwidth supported by all STAs in the BSS. The bandwidth of the primary channel can be set and / or limited by the STAs operating in the BSS that support the minimum bandwidth operating mode. In the 802.11ah example, for STAs that support (e.g., only support) the 1 MHz mode (e.g., MTC type devices), the primary channel can be 1 MHz wide, even if the AP and other STAs in the BSS support 2 MHz, 4 MHz, 8 MHz, 16 MHz, and / or other channel bandwidth operating modes. Carrier Sense and / or Network Assignment Vector (NAV) settings can depend on the status of the primary channel. For example, if the primary channel is busy due to STAs (which only support the 1 MHz operating mode) transmitting to the AP, the entire available band can be considered busy, even if most of the band remains idle and may be available.
[0065] In the United States, the available frequency band for 802.11ah is 902 MHz to 928 MHz. In South Korea, the available frequency band is 917.5 MHz to 923.5 MHz. In Japan, the available frequency band is 916.5 MHz to 927.5 MHz. The total available bandwidth for 802.11ah is 6 MHz to 26 MHz, depending on the country code.
[0066] Figure 1D This is a system diagram illustrating RAN 113 and CN 115 according to an embodiment. As described above, RAN 113 can communicate with WTRUs 102a, 102b, and 102c via air interface 116 using NR radio technology. RAN 113 can also communicate with CN 115.
[0067] RAN 113 may include gNBs 180a, 180b, and 180c; however, it should be understood that RAN 113 may include any number of gNBs while remaining consistent with the embodiments. gNBs 180a, 180b, and 180c may each include one or more transceivers for communicating with WTRUs 102a, 102b, and 102c via air interface 116. In one embodiment, gNBs 180a, 180b, and 180c may implement MIMO technology. For example, gNBs 180a and 180b may utilize beamforming to transmit signals to and / or receive signals from gNBs 180a, 180b, and 180c. Thus, for example, gNB 180a may use multiple antennas to transmit radio signals to and / or receive radio signals from WTRU 102a. In embodiments, gNBs 180a, 180b, and 180c may implement carrier aggregation technology. For example, gNB 180a can transmit multiple component carriers (not shown) to WTRU 102a. A subset of these component carriers may be on unlicensed spectrum, while the remaining component carriers may be on licensed spectrum. In embodiments, gNBs 180a, 180b, and 180c may implement Coordinated Multipoint (CoMP) technology. For example, WTRU 102a may receive coordinated transmissions from gNBs 180a and 180b (and / or gNB 180c).
[0068] WTRUs 102a, 102b, and 102c can communicate with gNBs 180a, 180b, and 180c using transmissions associated with a scalable set of parameters. For example, OFDM symbol spacing and / or OFDM subcarrier spacing can vary for different transmissions, different cells, and / or different portions of the radio transmission spectrum. WTRUs 102a, 102b, and 102c can communicate with gNBs 180a, 180b, and 180c using multiple or scalable length subframes or transmission time intervals (TTIs) (e.g., containing a variable number of OFDM symbols and / or a continuously variable length absolute time).
[0069] gNBs 180a, 180b, and 180c can be configured to communicate with WTRUs 102a, 102b, and 102c in standalone and / or non-standalone configurations. In standalone configuration, WTRUs 102a, 102b, and 102c can communicate with gNBs 180a, 180b, and 180c without additional access to other RANs (e.g., eNode-Bs 160a, 160b, and 160c). In standalone configuration, WTRUs 102a, 102b, and 102c can utilize one or more of gNBs 180a, 180b, and 180c as mobility anchors. In standalone configuration, WTRUs 102a, 102b, and 102c can communicate with gNBs 180a, 180b, and 180c using signals in unlicensed frequency bands. In a non-standalone configuration, WTRUs 102a, 102b, and 102c can communicate / connect with gNBs 180a, 180b, and 180c while also communicating / connecting with another RAN (such as eNode-Bs 160a, 160b, and 160c). For example, WTRUs 102a, 102b, and 102c can implement DC principles to communicate substantially simultaneously with one or more gNBs 180a, 180b, and 180c and one or more eNode-Bs 160a, 160b, and 160c. In a non-standalone configuration, eNode-Bs 160a, 160b, and 160c can act as mobility anchors for WTRUs 102a, 102b, and 102c, and gNBs 180a, 180b, and 180c can provide additional coverage and / or throughput for serving WTRUs 102a, 102b, and 102c.
[0070] Each of gNBs 180a, 180b, and 180c can be associated with a specific cell (not shown) and can be configured to handle radio resource management decisions, handover decisions, user scheduling in UL and / or DL, network slicing support, dual connectivity, interoperability between NR and E-UTRA, routing of user plane data to User Plane Functions (UPF) 184a and 184b, and routing of control plane information to Access and Mobility Management Functions (AMF) 182a and 182b, etc. Figure 1D As shown, gNB 180a, 180b, and 180c can communicate with each other via the Xn interface.
[0071] Figure 1DThe CN 115 shown may include at least one AMF 182a, 182b, at least one UPF 184a, 184b, at least one Session Management Function (SMF) 183a, 183b, and possibly a Data Network (DN) 185a, 185b. Although each of the foregoing elements is depicted as part of the CN 115, it should be understood that any of these elements may be owned and / or operated by an entity other than a CN operator.
[0072] AMF 182a and 182b can connect to one or more gNBs 180a, 180b, and 180c in RAN 113 via the N2 interface and can act as control nodes. For example, AMF 182a and 182b can be responsible for authenticating users of WTRU 102a, 102b, and 102c, supporting network slicing (e.g., handling different PDU sessions with different requirements), selecting specific SMF 183a and 183b, managing registration areas, terminating NAS signaling, mobility management, etc. AMF 182a and 182b can use network slicing to customize CN support for WTRU 102a, 102b, and 102c based on the service type used by WTRU 102a, 102b, and 102c. For example, different network slices can be established for different use cases, such as services relying on Ultra Reliable Low Latency (URLLC) access, services relying on Enhanced Massive Mobile Broadband (eMBB) access, and services for Machine Type Communication (MTC) access. The AMF162 can provide control plane functions for switching between RAN 113 and other RANs (not shown) that employ other radio technologies such as LTE, LTE-A, LTE-A Pro and / or non-3GPP access technologies such as WiFi.
[0073] SMF 183a and 183b can connect to AMF 182a and 182b in CN 115 via the N11 interface. SMF 183a and 183b can also connect to UPF 184a and 184b in CN 115 via the N4 interface. SMF 183a and 183b can select and control UPF 184a and 184b, and configure routing for service flows passing through UPF 184a and 184b. SMF 183a and 183b can perform other functions, such as managing and allocating UE IP addresses, managing PDU sessions, controlling policy enforcement and QoS, and providing downlink data notifications. PDU session types can be IP-based, non-IP-based, Ethernet-based, etc.
[0074] UPF 184a and 184b can connect to one or more gNBs 180a, 180b, and 180c in RAN 113 via the N3 interface. This N3 interface provides WTRU 102a, 102b, and 102c with access to packet-switched networks (such as Internet 110), facilitating communication between WTRU 102a, 102b, and 102c and IP-enabled devices. UPF 184 and 184b can perform other functions such as routing and forwarding packets, enforcing user plane policies, supporting multihomed PDU sessions, handling user plane QoS, buffering downlink packets, and providing mobility anchoring.
[0075] CN 115 can facilitate communication with other networks. For example, CN 115 may include or be able to communicate with an IP gateway (e.g., an IP Multimedia Subsystem (IMS) server) that acts as an interface between CN 115 and PSTN 108. Furthermore, CN 115 can provide WTRUs 102a, 102b, and 102c with access to other networks 112, which may include other wired and / or wireless networks owned and / or operated by other service providers. In one embodiment, WTRUs 102a, 102b, and 102c can be connected to local data networks (DNs) 185a and 185b via UPFs 184a and 184b through their N3 interfaces and the N6 interface between UPFs 184a and 184b and DNs 185a and 185b.
[0076] Given Figures 1A to 1D As described herein, one or more of the functions described herein may be performed by one or more emulation devices (not shown): WTRU 102a to 102d, base stations 114a and 114b, eNode-B 160a to 160c, MME 162, SGW 164, PGW 166, gNB 180a to 180c, AMF 182a and 182b, UPF 184a and 184b, SMF 183a and 183b, DN 185a and 185b, and / or any other devices described herein. An emulation device may be one or more devices configured to emulate one or more of the functions described herein. For example, an emulation device may be used to test other devices and / or simulate network and / or WTRU functions.
[0077] Simulation devices can be designed to perform one or more tests on other devices in laboratory and / or carrier network environments. For example, one or more simulation devices can perform one or more functions while being fully or partially implemented and / or deployed as part of a wired and / or wireless communication network to test other devices within the communication network. One or more simulation devices can perform one or more functions while being temporarily implemented / deployed as part of a wired and / or wireless communication network. Simulation devices can be directly coupled to another device for testing purposes and / or can perform tests using over-the-air wireless communication.
[0078] One or more simulation devices can perform one or more functions while not being implemented / deployed as part of a wired and / or wireless communication network. For example, simulation devices can be utilized in test scenarios within a test laboratory and / or an undeployed (e.g., testing) wired and / or wireless communication network to perform tests on one or more components. One or more simulation devices can be test equipment. Simulation devices can transmit and / or receive data using direct RF coupling and / or wireless communication via an RF circuit system (e.g., which may include one or more antennas).
[0079] This application describes various aspects, including tools, features, examples, models, methods, etc. Many of these aspects are described in detail, and are generally described in a way that may sound restrictive, at least to illustrate individual characteristics. However, this is merely for clarity and does not limit the application or scope of these aspects. In fact, all the different aspects can be combined and interchanged to provide further aspects. Furthermore, these aspects can be combined and interchanged with those described in earlier filings.
[0080] The aspects described and envisioned in this application can be implemented in many different forms. Figures 5 to 12 Some examples can be provided, but other examples are also envisioned. Figures 5 to 12 The discussion does not limit the breadth of implementations. At least one aspect relates generally to video encoding and decoding, and at least another aspect relates generally to the transmission of generated or encoded bitstreams. These and other aspects can be implemented as methods, apparatus, computer-readable storage media having instructions stored thereon for encoding or decoding video data according to any of the methods, and / or computer-readable storage media having bitstreams generated according to any of the methods stored thereon.
[0081] In this application, the terms “reconstruction” and “decoding” are used interchangeably, the terms “pixel” and “sample” are used interchangeably, and the terms “image”, “picture” and “frame” are used interchangeably.
[0082] This document describes various methods, and each method includes one or more steps or actions for implementing the method. Unless a specific order of steps or actions is required for the correct operation of the method, the order and / or use of specific steps and / or actions can be modified or combined. Furthermore, terms such as "first" and "second" can be used in various examples to modify elements, components, steps, operations, etc., for example, "first decoding" and "second decoding." Unless specifically required, the use of such terms does not imply the order of the modified operations. Therefore, in this example, the first decoding does not need to be performed before the second decoding and can occur, for example, before, during, or in a period overlapping with the second decoding.
[0083] The various methods and other aspects described in this application can be used to modify the module, for example, such as Figure 2 and Figure 3 The decoding modules of the video encoder 200 and decoder 300 are shown. Furthermore, the subject matter disclosed herein can be applied, for example, to: any type, format, or version of video encoding and decoding, whether described in standards or recommendations, whether pre-existing or future-developed; and any extensions to such standards and recommendations. Unless otherwise stated or technically excluded, the aspects described in this application may be used alone or in combination.
[0084] Various numerical values are used in the examples described in this application. These and other specific values are used for the purpose of describing the examples, and the aspects described are not limited to these specific values.
[0085] Figure 2 This is a diagram illustrating an exemplary video encoder. Variations of the exemplary encoder 200 are envisioned, but for clarity, encoder 200 is described below without describing all anticipated variations.
[0086] Before encoding, the video sequence can undergo pre-coding (201), for example, applying color transformations to the input color image (e.g., converting from RGB 4:4:4 to YCbCr 4:2:0), or performing remapping on the input image components to obtain a signal distribution more robust to compression (e.g., using histogram equalization with one of the color components). Metadata can be associated with pre-processing and appended to the bitstream.
[0087] In encoder 200, the image is encoded by encoder elements as described below. The image to be encoded is divided (202) and processed, for example, in units of coding units (CUs). Each unit is encoded, for example, using an intra-frame or inter-frame mode. When a unit is encoded in intra-frame mode, it performs intra-frame prediction (260). In inter-frame mode, motion estimation (275) and compensation (270) are performed. The encoder determines (205) which of the intra-frame or inter-frame modes to use to encode the unit, and indicates the intra-frame / inter-frame decision, for example, by a prediction mode flag. For example, the prediction residual (210) is calculated by subtracting the predicted block from the original image block.
[0088] Then, the predicted residual is transformed (225) and quantized (230). The quantized transform coefficients, along with the motion vector and other syntax elements, are entropy encoded (245) to output a bitstream. The encoder can skip the transform and apply quantization directly to the untransformed residual signal. The encoder can bypass both the transform and quantization, i.e., directly encode the residual without applying the transform or quantization process.
[0089] The encoder decodes the encoded blocks to provide a reference for further prediction. The quantized transform coefficients are dequantized (240) and inverse transformed (250) to decode the prediction residuals. The image blocks are reconstructed by combining (255) the decoded prediction residuals and the predicted blocks. A loop filter (265) is applied to the reconstructed image to perform, for example, deblocking / SAO (Sample Adaptive Offset) filtering to reduce coding artifacts. The filtered image is stored in a reference image buffer (280).
[0090] Figure 3 This is a diagram illustrating an exemplary video decoder. In the exemplary decoder 300, the bitstream is decoded by decoder elements as described below. The video decoder 300 typically performs the same operations as... Figure 2 The encoding process described herein is the inverse of the decoding process. Encoder 200 typically also performs video decoding as part of the video data encoding.
[0091] Specifically, the input to the decoder includes a video bitstream, which can be generated by the video encoder 200. First, entropy decoding (330) is performed on the bitstream to obtain transform coefficients, motion vectors, and other encoded information. Image partitioning information indicates how the image should be divided. The decoder can then segment (335) the image based on the decoded image partitioning information. The transform coefficients are dequantized (340) and inverse transformed (350) to decode the prediction residuals. Image blocks are reconstructed by combining (355) the decoded prediction residuals and the predicted blocks. The predicted blocks can be obtained from intra-frame prediction (360) or motion-compensated prediction (i.e., inter-frame prediction) (375). A loop filter (365) is applied to the reconstructed image. The filtered image is stored in a reference image buffer (380).
[0092] The decoded image can be further processed by post-decoding (385), such as inverse color transformation (e.g., from YCbCr 4:2:0 to RGB 4:4:4) or inverse remapping of the inverse of the remapping process performed in the pre-encoding process (201). Post-decoding can utilize metadata exported in the pre-encoding process and signaled in the bitstream. In the example, the decoded image (e.g., after applying a loop filter (365) and / or after post-decoding (385) if post-decoding is used) can be sent to a display device for presentation to the user.
[0093] Figure 4 This is a diagram illustrating exemplary systems that can implement the various aspects and examples described herein. System 400 can be embodied as a device including the various components described below and configured to perform one or more aspects described in this document. Examples of such devices include, but are not limited to, various electronic devices such as personal computers, laptop computers, smartphones, tablet computers, digital multimedia set-top boxes, digital television receivers, personal video recording systems, networked home appliances, and servers. Elements of system 400 (individually or in combination) can be embodied in a single integrated circuit (IC), multiple ICs, and / or discrete components. For example, in at least one example, the processing and encoder / decoder elements of system 400 are distributed across multiple ICs and / or discrete components. In many examples, system 400 is communicatively coupled to one or more other systems or other electronic devices, for example, via a communication bus or through dedicated input and / or output ports. In many examples, system 400 is configured to implement one or more aspects described in this document.
[0094] System 400 includes at least one processor 410 configured to execute instructions loaded thereon for implementing various aspects, such as those described in this document. Processor 410 may include embedded memory, input / output interfaces, and various other circuit systems known in the art. System 400 includes at least one memory 420 (e.g., a volatile memory device and / or a non-volatile memory device). System 400 includes a storage device 440, which may include non-volatile memory and / or volatile memory, including but not limited to electrically erasable programmable read-only memory (EEPROM), read-only memory (ROM), programmable read-only memory (PROM), random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), flash memory, disk drives, and / or optical disk drives. Storage device 440 may include internal storage devices, attached storage devices (including removable and non-removable storage devices), and / or network-accessible storage devices, as non-limiting examples.
[0095] System 400 includes an encoder / decoder module 430 configured to, for example, process data to provide encoded or decoded video, and the encoder / decoder module 430 may include its own processor and memory. The encoder / decoder module 430 represents a module that can be included in a device to perform encoding and / or decoding functions. It is well known that a device can include one or both encoding and decoding modules. Furthermore, the encoder / decoder module 430 may be implemented as a separate element of system 400, or it may be incorporated into processor 410 as a combination of hardware and software, as is known to those skilled in the art.
[0096] Program code to be loaded onto processor 410 or encoder / decoder 430 to execute the various aspects described in this document may be stored in storage device 440 and subsequently loaded onto memory 420 for execution by processor 410. According to various examples, one or more of processor 410, memory 420, storage device 440, and encoder / decoder module 430 may store one or more items during the execution of the processes described in this document. Such stored items may include, but are not limited to, input video, decoded video or a portion thereof, bitstreams, matrices, variables, and intermediate or final results from processing of equations, formulas, operations, and operational logic.
[0097] In some examples, the memory within processor 410 and / or encoder / decoder 430 is used to store instructions and provide working memory for the processing required during encoding or decoding. However, in other examples, external memory (e.g., the processing device may be processor 410 or encoder / decoder module 430) is used for one or more of these functions. External memory may be, for example, memory 420 and / or storage device 440, such as volatile memory and / or non-volatile flash memory. In several examples, external non-volatile flash memory is used to store, for example, the operating system of a television. In at least one example, fast external volatile memory (such as RAM) is used as working memory for video encoding and decoding operations.
[0098] Inputs to the components of system 400 can be provided by a variety of input devices as shown in box 445. Such input devices include, but are not limited to, (i) a radio frequency (RF) section that receives, for example, RF signals transmitted over the air by a broadcaster, (ii) component input terminals (or sets of COMP input terminals), (iii) universal serial bus (USB) input terminals, and / or (iv) high-definition multimedia interface (HDMI) input terminals. Figure 4 Other examples not shown include composite video.
[0099] In various examples, the input device of block 445 has corresponding associated input processing elements known in the art. For example, the RF section may be associated with elements suitable for: (i) selecting a desired frequency (also known as selecting a signal, or band-limiting a signal to a band), (ii) down-converting the selected signal, (iii) band-limiting it again to a narrower band to select, for example, a signal band (which may be referred to as a channel in some examples), (iv) demodulating the down-converted and band-limited signal, (v) performing error correction, and / or (vi) demultiplexing to select a desired data packet stream. The RF section of various examples includes one or more elements for performing these functions, such as a frequency selector, signal selector, band limiter, channel selector, filter, downconverter, demodulator, error corrector, and demultiplexer. The RF section may include a tuner that performs multiple of these functions, including, for example, down-converting a received signal to a lower frequency (e.g., intermediate frequency or near-baseband frequency) or baseband. In one set-top box example, the RF section and its associated input processing elements receive RF signals transmitted via a wired (e.g., cable) medium and perform frequency selection by filtering, down-converting, and re-filtering to the desired frequency band. Various examples rearrange the order of the above (and other) components, remove some of these components, and / or add other components that perform similar or different functions. Adding components may include inserting components between existing components, such as inserting amplifiers and analog-to-digital converters. In many examples, the RF section includes an antenna.
[0100] USB and / or HDMI terminals may include corresponding interface processors for connecting system 400 to other electronic devices across USB and / or HDMI connections. It should be understood that multiple aspects of input processing (e.g., Reed-Solomon error correction) may be implemented as needed, for example, within a separate input processing IC or within processor 410. Similarly, aspects of USB or HDMI interface processing may be implemented as needed, either within a separate interface IC or within processor 410. The demodulated, error-corrected, and demultiplexed stream is provided to various processing elements, including, for example, processor 410 and encoder / decoder 430, which operate in conjunction with memory and storage elements to process the data stream as needed for presentation on the output device.
[0101] Various components of system 400 can be housed within an integrated housing. Within the integrated housing, various components can be interconnected and transmit data to each other using suitable connection means 425 (e.g., internal buses known in the art, including inter-IC (I2C) buses, wiring, and printed circuit boards).
[0102] System 400 includes a communication interface 450 that enables communication with other devices via a communication channel 460. The communication interface 450 may include, but is not limited to, a transceiver configured to send and receive data via the communication channel 460. The communication interface 450 may include, but is not limited to, a modem or network interface card (NIC), and the communication channel 460 may be implemented, for example, within a wired and / or wireless medium.
[0103] In several examples, data is provided to system 400 via streaming or otherwise using a wireless network, such as a Wi-Fi network, such as IEEE 802.11 (IEEE stands for Institute of Electrical and Electronics Engineers). In these examples, the Wi-Fi signal is received via a communication channel 460 and a communication interface 450 adapted for Wi-Fi communication. The communication channel 460 in these examples is typically connected to an access point or router that provides access to external networks, including the Internet, to allow streaming applications and other over-the-top communications. Other examples use a set-top box that provides data via an HDMI connection in input box 445 to provide streaming data to system 400. Still other examples use an RF connection in input box 445 to provide streaming data to system 400. As mentioned above, several examples provide data in a non-streaming manner. Furthermore, several examples use wireless networks other than Wi-Fi, such as cellular networks or Bluetooth® networks.
[0104] System 400 can provide output signals to a variety of output devices, including a display 475, a speaker 485, and other peripheral devices 495. Display 475 may include one or more of, for example, a touchscreen display, an organic light-emitting diode (OLED) display, a curved display, and / or a foldable display. Display 475 may be a display for a television, tablet computer, laptop computer, cellular phone (mobile phone), or other device. Display 475 may also be integrated with other components (e.g., as in a smartphone) or separate (e.g., an external monitor for a laptop computer). Other peripheral devices 495 may include one or more of, in various examples, a standalone digital video disc (or digital multifunction disc) (DVD, for both terms), an optical disc player, a stereo system, and / or a lighting system. Various examples utilize one or more peripheral devices 495 that provide functionality based on the output of system 400. For example, an optical disc player performs the function of playing the output of system 400.
[0105] In various examples, control signals communicate between system 400 and display 475, speaker 485, or other peripheral devices 495 using signaling such as AV.Link, Consumer Electronics Control (CEC), or other support device-to-device control (with or without user intervention) communication protocols. Output devices can be communicatively coupled to system 400 via dedicated connections through corresponding interfaces 470, 480, and 490. Alternatively, output devices can be connected to system 400 via communication interface 450 using communication channel 460. Display 475 and speaker 485 can be integrated into a single unit with other components of system 400 in an electronic device such as a television. In various examples, display interface 470 includes a display driver, such as, for example, a timing controller (TCon) chip.
[0106] Display 475 and speaker 485 may alternatively be separate from one or more other components, for example, if the RF section of input 445 is part of a separate set-top box. In various examples where display 475 and speaker 485 are external components, the output signal may be provided via a dedicated output connection, including, for example, an HDMI port, a USB port, or a COMP output.
[0107] These examples can be implemented by computer software or hardware, or a combination of hardware and software, implemented by processor 410. As a non-limiting example, these examples can be implemented by one or more integrated circuits. Memory 420 can be of any type suitable for the technical environment and can be implemented using any suitable data storage technology, such as optical memory devices, magnetic memory devices, semiconductor-based memory devices, fixed memory, and removable memory, as non-limiting examples. Processor 410 can be of any type suitable for the technical environment and can encompass one or more microprocessors, general-purpose computers, special-purpose computers, and processors based on multi-core architectures, as non-limiting examples.
[0108] Various implementations involve decoding. As used herein, “decoding” can encompass, for example, all or part of a process performed on a received encoded sequence to produce a final output suitable for display. In various examples, such a process includes one or more processes typically performed by a decoder, such as entropy decoding, inverse quantization, inverse transform, and differential decoding. In various examples, such a process further or alternatively includes processes performed by a decoder of various embodiments described herein, such as decoding video data to determine partitioning information, determining partitions, partition groups, and group coding technique information associated with each partition group, determining a decoding technique for the partition group based on the partitioning information, determining signal values for the partition group based on the determined decoding technique, and reconstructing the signal based on the determined signal values, etc.
[0109] As further examples, in one example, "decoding" refers only to entropy decoding; in another example, "decoding" refers only to differential decoding; and in yet another example, "decoding" refers to a combination of entropy decoding and differential decoding. Whether the phrase "decoding process" is intended to specifically refer to a subset of operations or to refer to a broader decoding process will be clear based on the specific context of the description and is thought to be well understood by those skilled in the art.
[0110] Various implementations involve encoding. In a manner similar to the above discussion of “decoding,” the term “encoding,” as used herein, can encompass, for example, all or part of the process performed on an input video sequence to produce an encoded bitstream. In many examples, such a process includes one or more processes typically performed by an encoder, such as partitioning, differential coding, transform, quantization, and entropy coding. In many examples, such a process further or alternatively includes processes performed by an encoder of various embodiments described herein, such as determining partitions of the input signal, determining partition groups, determining representation parameters for each partition group, encoding the partitions, partition groups, and representation parameters, including encoded information in the video data, etc.
[0111] As further examples, in one example, "encoding" refers only to entropy encoding; in another example, "encoding" refers only to differential encoding; and in yet another example, "encoding" refers to a combination of differential and entropy encoding. Whether the phrase "encoding process" is intended to specifically refer to a subset of operations or to refer to a broader encoding process will be clear based on the specific context of the description and is thought to be well understood by those skilled in the art.
[0112] It should be noted that the syntax elements used in this article (such as the encoding syntax for coding unit groups, INR parameters, etc.) are descriptive terms. Therefore, the use of other syntax element names is not excluded.
[0113] When a diagram is presented as a flowchart, it should be understood that it also provides a block diagram of the corresponding device. Similarly, when a diagram is presented as a block diagram, it should be understood that it also provides a flowchart of the corresponding method / process.
[0114] The embodiments and aspects described herein can be implemented, for example, in methods or processes, apparatuses, software programs, data streams, or signals. Even if discussed only in the context of a single embodiment (e.g., discussed only as a method), embodiments of the discussed features can be implemented in other forms (e.g., apparatuses or programs). Apparatuses can be implemented, for example, in suitable hardware, software, and firmware. Methods can be implemented, for example, in a processor, which generally refers to a processing device, including, for example, a computer, microprocessor, integrated circuit, or programmable logic device. Processors also include communication devices, such as, for example, computers, cellular phones, portable / personal digital assistants (“PDAs”), and other devices that facilitate information communication between end users.
[0115] References to “an example” or “an example” or “an implementation” or “an implementation” and other variations mean that a particular feature, structure, characteristic, etc., described in connection with that example is included in at least one example. Therefore, the phrases “in an example” or “in an example” or “in an implementation” or “in an implementation” and any other variations appearing in various places in this application do not necessarily refer to the same example.
[0116] Furthermore, this application may refer to "determining" various types of information. Determining information may include one or more of, for example, estimation information, calculation information, prediction information, or information retrieved from memory. Obtaining may include receiving, retrieving, constructing, generating, and / or determining.
[0117] Furthermore, this application may refer to "accessing" various types of information. Accessing information may include, for example, receiving information, retrieving information (e.g., from memory), storing information, moving information, copying information, calculating information, determining information, predicting information, or estimating information, one or more of these.
[0118] Furthermore, this application may refer to "receiving" various types of information. Like "accessing," "receiving" is intended to be a broad term. Receiving information may include, for example, accessing information or retrieving information (e.g., from memory) one or more of these. Moreover, "receiving" is generally involved in some way during operations such as storing information, processing information, sending information, moving information, copying information, erasing information, calculating information, determining information, predicting information, or estimating information.
[0119] It should be understood that the use of any of the following “ / ”, “and / or”, and “…at least one of…”, such as in the cases of “A / B”, “A and / or B”, and “at least one of A and B”, is intended to cover selecting only the first listed option (A), or only the second listed option (B), or both options (A and B). As a further example, in the cases of “A, B, and / or C” and “at least one of A, B, and C”, such wording is intended to cover selecting only the first listed option (A), or only the second listed option (B), or only the third listed option (C), or only the first and second listed options (A and B), or only the first and third listed options (A and C), or only the second and third listed options (B and C), or all three options (A, B, and C). As will be apparent to those skilled in the art and related fields, this can be extended to as many items as possible listed.
[0120] Furthermore, as used herein, the term “signal” refers, among other things, to instructing the corresponding decoder to do something. Encoder signals may include, for example, partitions, partition groups, representation parameters, encoding techniques, etc. In this way, in the examples, the same parameters are used on both the encoder and decoder sides. Thus, for example, the encoder may (explicitly signal) send specific parameters to the decoder so that the decoder can use the same specific parameters. Conversely, if the decoder already has specific parameters as well as other parameters, it may use signaling without transmission (implicit signaling) to simply allow the decoder to know and select specific parameters. Bit savings are achieved in various examples by avoiding the transmission of any actual functions. It should be understood that signaling can be done in various ways. For example, in various examples, one or more syntax elements, flags, etc., are used to signal information to the corresponding decoder. Although the foregoing refers to the verb form of the term “signal,” the word “signal” can also be used as a noun in this document.
[0121] As will be apparent to those skilled in the art, implementations can generate various signals formatted to carry, for example, information that can be stored or transmitted. This information may include, for example, instructions for performing a method, or data generated by one of the described implementations. For example, the signal may be formatted to carry a bitstream of the described example. Such a signal may be formatted, for example, as an electromagnetic wave (e.g., using the radio frequency portion of the spectrum) or a baseband signal. Formatting may include, for example, encoding a data stream and modulating a carrier wave with the encoded data stream. The information carried by the signal may be, for example, analog or digital information. The signal may be transmitted via a variety of different wired or wireless links as known. The signal may be stored on, accessed from, or received from a processor-readable medium.
[0122] This document describes numerous examples. Features of the examples may be provided individually or in any combination across multiple claim classes and types. Furthermore, examples may include one or more features, devices, or aspects described herein, individually or in any combination across multiple claim classes and types. For example, features described herein may be implemented in a bitstream or signal including information generated as described herein. This information may allow a decoder to decode the bitstream, encoder, bitstream, and / or decoder according to any of the described embodiments. For example, features described herein may be implemented by creating and / or transmitting and / or receiving and / or decoding a bitstream or signal. For example, features described herein may be implemented as a method, process, apparatus, medium storing instructions, medium storing data, or signal. For example, features described herein may be implemented by a television, set-top box, cellular phone, tablet computer, or other electronic device performing decoding. The television, set-top box, cellular phone, tablet computer, or other electronic device may display (e.g., using a monitor, screen, or other type of display) a resulting image (e.g., an image reconstructed from the residual of a video bitstream). The television, set-top box, cellular phone, tablet computer, or other electronic device may receive a signal including an encoded image and perform decoding.
[0123] Systems, methods, and tools for performing superpixel spatial merging for implicit neural representations (INRs) are disclosed. The signal can be encoded, for example, using partitions and INRs. Merging or grouping of portions of the signal can be supported and / or permitted. Information from the merging of the portions of the signal can be used to train representation parameters (e.g., INR parameters) for the grouped portions of the signal. This information can be encoded in video data (e.g., a bitstream). The video data (e.g., encoded video data) can be decoded to determine this information.
[0124] Devices (e.g., video encoders or video decoders) can be configured to perform the partitioning and merging of signals. The device can acquire an input signal. The device can determine partitions for the input signal. The device can determine groupings of the determined partitions. The device can determine appropriate representation parameters for each partition group. The device can determine the encoding technique associated with the partition group. The device can encode the partitions, partition groups, and representation parameters associated with each partition group. The device can include the encoded partitions, partition groups, and representation parameters associated with each partition group in the video data. The video data can include the determined encoding technique associated with the partition group. The video data can include instructions indicating the use of coding unit groups.
[0125] Partition groups can be determined. For example, partition groups can be determined based on the distance between partitions. Partition groups can be determined based on the impact of grouping on coding performance (e.g., based on using greedy search, brute force methods, reinforcement learning, genetic algorithms, etc.). Partition groups can be determined based on a trained model (e.g., a machine learning model). The device can train the model to generate partition grouping predictions.
[0126] Devices (e.g., video encoding or decoding devices) can decode video data to determine partitioning information. Partitioning information may indicate partitions, partition groups, and corresponding group coding technology information associated with each partition group. The device can determine the decoding technology for the partition group based on the partitioning information. The device can determine the signal values for the partition group based on the determined decoding technology. The device can reconstruct the signal, for example, based on the signal values.
[0127] The systems, methods, and tools described herein may relate to decoders. In some examples, the systems, methods, and tools described herein may relate to encoders. In some examples, the systems, methods, and tools described herein may relate to signals (e.g., from an encoder and / or received by a decoder). Computer-readable media may include instructions for causing one or more processors to perform the methods described herein. A computer program product may include instructions that, when executed by one or more processors, cause one or more processors to perform the methods described herein.
[0128] The SympAI project can include superpixel spatial merging for implicit neural representations. Methods for neural compression can be developed. Implicit neural representations (INRs) can be used for image and / or video coding. INR-based compression techniques can be used. INRs can be investigated for use in 2D, video compression, and / or many other signals (e.g., 3D scenes or objects). These methods can use (e.g., have) (e.g., significantly) lower computational complexity compared to end-to-end neural compression methods.
[0129] Figure 5 The illustration shows an exemplary simple neural network used for implicit neural representation (INR). Such a neural network used for INR can be called an INR network. INR can be used for signals of any dimension (e.g., any dimensional signal). Figure 5 An exemplary illustration of a 2D signal (such as an image) is shown in the image. INR can be used to parameterize a signal as a function (e.g., as...). Figure 5 (as shown in 500), this function will take the coordinates (e.g., as shown in 500). Figure 5 (as shown in 510) serves as a potential approximation of the input and output signal at these coordinates (e.g., as shown in 510). Figure 5(As shown in 520). INR can be applied to images, 2D video, or 3D objects, as well as other applications. In the case of images, the input (e.g., as shown in 520) is... Figure 5 As shown in 510, it can include pixel coordinates (x, y), and the INR output (e.g., as shown in 510) ... Figure 5 As shown in 520, the output can include the color values (r, g, b) or (y, u, v) of the input pixels. In the case of video, the output can be similar, and the input can include, for example, the frame index t in addition to the pixel coordinates.
[0130] Inter-noise refraction (INR) can be used to reconstruct signals by computing signal values for coordinate inputs. The input coordinates can be modified, for example, by a transformation before being used as input to a neural network. This transformation can be a Fourier map, coordinate transformation, normalization, etc.
[0131] INR networks (e.g., Figure 5 The 500 shown in the diagram can (e.g., typically) be a neural network, which consists of multiple neural layers (e.g., fully connected layers). Figure 5 As shown, the network can include four layers. Intermediate outputs can be represented by circles. Each neural layer can be described as a function that can (e.g., first) multiply the input by a tensor, add a vector called a bias, and / or (e.g., then) apply a nonlinear function to the resulting value. The shape and / or other properties of the tensor, as well as the type of the nonlinear function, can be called the network's architecture. The values of the tensors and the biases can be represented using the term weights. The parameters of the weights and the nonlinear function (e.g., if applicable) can be called the network's parameters θ. The architecture and parameters define the model. The function f_θ can represent an INR function parameterized by θ.
[0132] Figure 6 The diagram illustrates the use of INR to encode signals (e.g., as shown in the image). Figure 6 The exemplary process is shown in 610 (in the example). This can be achieved by optimizing the weights θ (e.g., or a subset of weights) of the INR network to reconstruct the signal (e.g., as shown in 610). Figure 6 This can be accomplished as shown in 620 (in the example). In some examples, the weights can be encoded (e.g., as shown in 620). Figure 6 As shown in 630) to create output video data (e.g., bitstream) (e.g., as shown in 630) Figure 6 (As shown in 650). For an image I of size (M×N), the weights θ can be optimized, for example, by minimizing the following loss function according to equations 1 and 2: Equation 1 Equation 2 Where D can indicate distortion, and its quantization is determined by f θThe difference between the reconstructed image and the original image I, R can indicate the bit rate of the encoded parameters, and λ can be a trade-off parameter between D and R. D can be (e.g., any) differentiable distortion metric, such as the mean squared error in the second equation. M and N can be the width and height of the image. Other metrics (such as learned perceptual image patch similarity (LPIPS)) can also be used. The optimization of the weights θ can be (e.g., typically) performed by machine learning methods (e.g., batch gradient descent).
[0133] To decompress the signal, f can be evaluated at the relevant coordinates. θ These coordinates can be selected (e.g., during decoding). A typical selection could include pixel coordinates for an image or video. In an example (e.g., for a 256x256 pixel image), these coordinates could be (e.g., all) pairs of (x, y), where (e.g., all) x ∈ {0,1,...,255} and y ∈ {0,1,...,255}. Other selections are also possible, such as for upsampling, downsampling, or expanding the original image.
[0134] In some examples, the input coordinates can be preprocessed, for instance, before being used as input to an implicit neural network. Preprocessing can involve one or more of the following operations: coordinate normalization between 0 and 1, between -1 and 1, or other values; transformation to another coordinate system (e.g., polar or Euclidean coordinates); mapping to other values (e.g., using a Fourier mapping), and so on. The Fourier mapping of pixel coordinates v=(x,y) can be defined accordingly according to Equation 3: Equation 3 The mapping can depend on the coefficient a i b i For example, where the coefficient b is considered as a Fourier approximation of the kernel function if (e.g., when) the mapping is viewed as such. i These can be Fourier basis frequencies. The preprocessed coordinates can be represented by p(x,y).
[0135] In some examples, using a (e.g., a single) INR network on the entire signal can (e.g., often) be suboptimal. The domain of the input signal can be divided into connected parts (e.g., coding units (CUs)). This partitioning can take any form, such as fixed-size partitions, superpixels, coding trees, etc. Figure 7 The illustration depicts an exemplary encoding process. First, the input signal (e.g., such as...) is... Figure 7 (as shown in 710) division (e.g., as shown in 710) Figure 7(As shown in 720). The signal can be (e.g., then) encoded using INR representations (e.g., a set of INR representations). A bijection can exist between the set of INR representations and a portion of the signal domain. Each portion of the signal domain can be encoded by its associated INR representation, and the associated INR representation can be trained on that portion of the signal (e.g., as shown in 720). Figure 7 (As shown in 730). The signal can be encoded by encoding the partitions (e.g., if needed) and the parameters represented by different INRs (e.g., as shown in 730). Figure 7 (As shown in 740) into video data (e.g., bitstream).
[0136] The encoded signal may include parameters for optimizing (e.g., each) the INR network. The weights can be optimized, for example, by a loss function involving the following distortion metric according to Equation 4: Equation 4 Where L can indicate the number of parts of the signal domain, C i It can indicate the portion of the signal domain indexed by i, and θ i Is with C i Associated parameters. Symbol p Ci It can be emphasized (e.g., some) that preprocessing operations may depend on C. i In the example, normalization can be performed relative to the values of the coordinates in that section (e.g., only). The rate can be adjusted (e.g., in the resulting loss function) to measure the bit length of the parameter encoding.
[0137] Multiple parts of a signal can be encoded together. For example, one object may (e.g., partially) occlude another object, so the occluded object may be covered by discontinuous portions of the signal. Encoding these parts of the signal together can achieve better coding efficiency, for example, by reusing (e.g., some or all) parameters of the INR or by encoding these parts together.
[0138] In the context of dividing a signal (e.g., a 2D image) and encoding it using a set of INRs, one of the following may be used, supported, and / or provided: merging or grouping parts of the signal (e.g., including discontinuous parts); encoding methods that utilize this information when training INR parameters for grouped parts of the signal; mechanisms for encoding these information fragments in video data (e.g., bitstreams); and / or associated decoding processes, etc.
[0139] Superpixels can be used as part of a signal (e.g., and / or other use cases).
[0140] It may support and / or allow the encoder to group or combine portions of a signal and process them together for encoding (e.g., to further improve encoding), for example, in the context of encoding a signal (e.g., 2D video) by partitioning it and using a set of INRs. Such grouping or merging of signal portions can be referred to as part groups. Certain portions of the signal (e.g., even if they are not connected or have been partitioned separately) may be similar and may share (e.g., some or all) parameters of the INR. For example, some objects or textures in the background may be occluded and segmented by another object in the foreground (e.g., partially). Portions of the background objects can be encoded together.
[0141] Figure 8 The diagram illustrates a pair of inputs (e.g., such as) used for sequential encoding. Figure 8 An exemplary encoding process (shown in 810) is described. Sequential encoding may include one or more of the following.
[0142] Signal domain partitioning is possible (e.g., ... Figure 8 (As shown in 820). The signal can be taken as input, and the output can include partitions of the signal domain. Different partitioning methods can be used / selected (e.g., many off-the-shelf methods are available for partitioning). Selection can include the type of part considered. The components or units of the signal part resulting from partitioning can be called CUs. Possible choices can include quadtrees, superpixels, binary trees, ternary trees, multi-type trees, etc. Selection can include the optimization algorithm used. A brute-force approach can be used (e.g., considering possible partitions). One approach can include greedy search, where an initial coding unit (CU) covering the entire signal can be partitioned (e.g., incrementally) by evaluating the impact of CU partitioning. For example, if the impact in terms of rate / distortion is positive, the partitioning can be performed (e.g., completed). If the CU encoding (e.g., along with its INR) is less advantageous in terms of rate / distortion than encoding its sub-blocks, the CU can be further partitioned. CU partitioning can be constructed, for example, without the INR network being learned. For example, brute-force or greedy search methods can both optimize other properties of the CU, such as the pixel mean, variance, texture, and / or any other statistics of the signal within the considered CUs. A model that has already been trained to output the input signal (e.g., a machine learning model) can be used.
[0143] Partitions (e.g., sections) can be grouped (e.g., such as...) Figure 8 (As shown in 830). Grouping of portions can involve creating groups of portions of the signal. Objectives may include creating groups of portions that can subsequently lead to more efficient coding in the process. Several methods may be used (e.g., there may be) to group portions.
[0144] One example approach could include calculating the distances between parts and (e.g., then) grouping the parts based on those distances.
[0145] The distance can be (e.g., any) distance between subsets of coordinates belonging to each CU, for example, including one or more of the following: the shortest distance between any element of one CU and any element of another CU or the distance between the centroids of each CU; distances that measure the similarity of signal values in these CUs, such as (e.g., any) Wasserstein distances on the distributions of these signal values, the average of the Wasserstein distances on the distributions of these signal values in each dimension, the quadratic distances between the histograms of these signal values, or the KL divergence between the approximations of these signal values; possible weighted combinations of these distances; and / or similar measures.
[0146] CU groups can be created, for example, based on the calculated pairwise distances between CUs. Different methods can be used (e.g., possibly). For example, a threshold δ can be chosen. (E.g., then) CU groups can be created by iterating through the CUs and adding CUs that are not in a group and whose distance to the currently considered CU is less than δ to the group. A CU group can be defined as the largest group such that the pairwise distances between all CUs in the group are less than δ. CU groups can be created by sorting the pairwise distances less than δ in ascending order, iterating through these distances, and adding any CU that is not already in a group to another group. Additional constraints can also be used, such as group size limits, constraints on group connectivity, constraints on the ratio between the largest and smallest CUs, etc.
[0147] One example approach could include directly considering the impact of grouping on coding performance and optimizing these groupings using any optimization algorithm (e.g., greedy search, brute force, reinforcement learning, genetic algorithm, etc.).
[0148] One example approach could include training a machine learning model to predict whether the two CUs should be grouped.
[0149] The INR network can be trained (e.g., as shown in the image). Figure 8 (As shown in 840). An INR can be trained (e.g., one) for each (e.g., each) CU group (e.g., independently). An INR can be trained for each CU that is not in a group (e.g., each). Learning INR parameters for CUs that do not belong to a group can be the same as or similar to learning a network for CUs that do not use a group. The CU itself can be considered as a signal.
[0150] INR network weights can be learned for CU groups to take advantage of the similarity between CUs. Options may include one or more of the following.
[0151] CUs in a CU group can be encoded together as if they were a single CU. In this case, preprocessing steps can be shared among all CUs.
[0152] CUs can share some parameters of the INR, such as some layers of the INR. This can include selecting shared parameters and learning shared and non-shared parameters. Parameter selection can include evaluating the performance of different possible choices on group encoding, comparing these performances with other choices, for example, selecting a set of parameters based on the characteristics of the group and individual CUs (e.g., directly) through an expert-designed algorithm or machine learning model. Parameter learning can then be performed, for example, by jointly learning parameters for the CUs in the group, or by learning INR parameters for a subset of CUs in the group and then reusing the values of the shared parameters for the INR of the other CUs in the group.
[0153] Encoding can be performed (e.g., such as...) Figure 8 (As shown in 850). Signal partitions, groups, group-level encoding techniques, and associated INR parameters can be encoded to create a bitstream for transmission or later use. This may involve using an entropy encoder. Signal partitions and INR parameters can be encoded. INR parameters can be encoded using exemplary codecs (e.g., MPEG-NNC). For groups, several encoding methods are possible.
[0154] Indications (e.g., flags) can be included in the video data (e.g., bitstream) to indicate the use of CU groups and their potential use in encoding.
[0155] Groups can be encoded using indices, where values are reserved for CUs that are not in the group. For example, a bitstream can contain (e.g., entropy-coded) index vectors, where the length of the vector can be the number of CUs.
[0156] Groups can be encoded by associating indices with CUs and by listing the indices of the CUs in the group. The size of a group can be signaled by reserving (e.g., one) an index value for the end of the group or by explicitly including the size of the group in the bitstream.
[0157] Group-level encoding techniques can be standardized and known to both the encoder and decoder. The encoder can (e.g., may be allowed) select the encoding technique. For example, the selected technique can be included in the bitstream and (e.g., potentially) entropy encoding can be performed. This information can be indicated (e.g., included) at the signal level, group level, or for a set number of groups. For some techniques, additional encoding may be required. Descriptions of shared parameters can be included in the bitstream, for example, if (e.g., when) parameters are shared between CUs in a group.
[0158] In some examples, a sequential approach can be used to optimize grouping and encoding. In some examples, several elements can be optimized simultaneously. Partitioning, grouping, parameters, and / or encoding length can be optimized together using any optimization algorithm, such as greedy search, gradient descent with a specific loss, genetic algorithms, the use of machine learning algorithms, etc. For example, grouping CUs can be based on the results of previous group encodings, for example, using reinforcement learning or active learning algorithms. For example, one or more steps may also involve computations that can benefit or be reused in another step. Some methods for partitioning signals or forming groups can use the calculation and evaluation of an INR for each possible part to assess the quality of such partitioning and / or grouping. In this case, these computations can be reused in another step, or the result of that other step can be obtained directly from the computation, or the two steps can be combined.
[0159] The resulting bitstream (e.g., such as Figure 8 As shown in 860, it may (e.g., in the instruction) include partitions, groups, encoding techniques for (e.g., each) the group, and INR parameters.
[0160] Such a bitstream can be decoded, for example, to obtain a reconstructed signal (e.g., as shown in the image). Figure 9 (As shown in 960). Figure 9 An exemplary process for generating the reconstructed signal is illustrated.
[0161] From the input bitstream (e.g., as Figure 9 As shown in 910), it can decode partitions, groups, encoding techniques, and parameter sets (e.g., such as...). Figure 9 (As shown in 920). This could involve using an entropy decoder and a dequantization operation.
[0162] For (e.g., each) CU group, it can be based, for example, on the coding technique used for that group and the associated parameters (e.g., such as...). Figure 9 (as shown in 930) to select the decoding technique, and / or the signal values for these CUs can be calculated, for example, based on an appropriate decoding technique (e.g., as shown in 930). Figure 9 (As shown in 940).
[0163] For example, if the group is encoded by encoding the CUs in the group together as if they were forming a single CU, the signal can be decoded as follows. For the input coordinates associated with the group (e.g., each), the signal values for those coordinates can be computed, for example, by performing inference with those coordinates as input using the reconstructed INR.
[0164] For example, if the group is encoded as having some parameters of the INR shared by all CUs in the group, then the group can be decoded as follows.
[0165] A network can be initialized (e.g., one) using shared parameter values. For (e.g., each) CU group C i The non-shared parameters of the network can be set to their values for the CU, and the signal values for the coordinates of the CU can be calculated by performing inference with these coordinates as input using INR.
[0166] In some examples, multiple groups can be decoded simultaneously. Multiple CUs can be decoded concurrently. Multiple networks can be built at once, and then values can be generated. Networks can be built, and the values of each CU can be computed sequentially, one at a time. The order of input coordinates within a CU can be modified. For example, batches of coordinates can be used as input to perform parallel computation of values.
[0167] Experimental results demonstrate that the method described in this paper leads to better rate / distortion performance (e.g., compared to not merging parts of the signal). The method can be applied to 24 images where the signal domain is divided into non-overlapping parts. These parts can then be merged based on the similarity of pixel colors within them. Each group of signal parts can then be encoded independently of other groups by training an INR. Figure 10 The illustration shows exemplary results achieved on an image with various λ values for the proposed method and baseline (where superpixels are not merged). Figure 10 The method (e.g., as described herein) achieves a BD-rate of -4.3% relative to the baseline. Bits per pixel can be reported without compressing network weights. These performance characteristics correspond to those achievable in intra-frame mode.
[0168] Figure 11 and 12 Examples of image partitioning for an image in the dataset before and after merging are provided respectively.
[0169] Figure 10 The illustration shows an example compression performance using the proposed invention (cross, red) relative to a baseline (where parts are not merged) (circles, blue). The proposed method (e.g., as described herein) can achieve a BD-rate of -4.3% relative to the baseline.
[0170] Figure 11 The illustration shows an example of superpixel segmentation before merging. Figure 12 The illustration shows an example of merged superpixel segmentation.
[0171] This method (such as that described in this paper) can be related to learning-based compression.
[0172] Furthermore, these methods can use (e.g., have) (e.g., far) lower computational complexity compared to end-to-end neural compression methods. INR methods can become localized (e.g., a set of local INRs) to achieve acceptable performance.
[0173] INR-based methods can enhance image / video compression using AI.
[0174] Although the features and elements have been described above in specific combinations, those skilled in the art will understand that each feature or element can be used alone or in any combination with other features and elements. Furthermore, the methods described herein can be embodied in a computer program, software, or firmware embodied in a computer-readable medium for execution by a computer or processor. Examples of computer-readable media include electronic signals (transmitted via wired or wireless connections) and computer-readable storage media. Examples of computer-readable storage media include, but are not limited to, read-only memory (ROM), random access memory (RAM), registers, cache memory, semiconductor memory devices, magnetic media such as internal hard disks and removable disks, magneto-optical media, and optical media such as CD-ROMs and digital versatile discs (DVDs). The processor associated with the software can be used to implement a radio frequency transceiver used in a WTRU, UE, terminal, base station, RNC, or any host computer.
Claims
1. A video encoding device, comprising: The processor is configured as follows: Determine a first partition associated with the input signal and a second partition associated with the input signal; Determine a first partition group and a second partition group, wherein the first partition group includes at least the first partition, and wherein the second partition group includes at least the second partition; Determine a first representation parameter associated with the first partition group and a second representation parameter associated with the second partition group; as well as The first partition, the second partition, the first partition group, the second partition group, the first representation parameter, and the second representation parameter are encoded. as well as The video data includes an encoded first partition, an encoded second partition, an encoded first partition group, an encoded second partition group, an encoded first representation parameter, and an encoded second representation parameter.
2. The video encoding device of claim 1, wherein the processor is further configured to: The distance between the first partition and the second partition is determined, wherein the first partition group and the second partition group are determined based on the distance between the first partition and the second partition.
3. The video encoding apparatus of any one of claims 1 and 2, wherein the processor is further configured to: The impact associated with partition groups is determined, wherein the impact associated with partition groups is determined based on the optimization of the coding performance associated with partition groups, wherein the first partition group and the second partition group are determined based on the impact associated with partition groups.
4. The video encoding apparatus of claim 3, wherein the optimization of encoding performance associated with partitioning and grouping includes using one or more of greedy search, brute force methods, reinforcement learning, and genetic algorithms.
5. The video encoding apparatus according to any one of claims 1 to 4, wherein the processor is further configured to: The model is trained to generate partition grouping predictions, wherein the first partition group and the second partition group are determined based on the trained model.
6. The video encoding apparatus of any one of claims 1 to 5, wherein the processor is further configured to: Determine a first coding technique associated with the first partition group and a second coding technique associated with the second partition group; and The video data includes the first encoding technique and the second encoding technique.
7. The video encoding apparatus of any one of claims 1 to 6, wherein the video data includes an indication to use a group of encoding units.
8. The video encoding apparatus of any one of claims 1 to 7, wherein the video data further includes a group information index, wherein the group information index indicates group information associated with the first partition and the second partition.
9. A video decoding device, comprising: The processor is configured as follows: Obtain the partition information associated with the video, wherein the partition information indicates partition group information; The decoding technique for the partition group is determined based on the partition information; The signal values associated with the partition group are determined based on the decoding technology. as well as The signal is reconstructed based on the determined signal values.
10. The video decoding apparatus of claim 9, wherein the processor is further configured to: Determine the distance between the first and second partitions; and The partition group is determined based on the distance between the first partition and the second partition.
11. The video decoding apparatus of any one of claims 9 and 10, wherein the processor is further configured to: Determine the impact associated with partition grouping, wherein the impact associated with partition grouping is determined based on optimizations of the coding performance associated with partition grouping; and The partition group is determined based on the effects associated with the partition grouping.
12. The video decoding apparatus of claim 11, wherein the optimization of the coding performance associated with partitioning and grouping includes using one or more of greedy search, brute force methods, reinforcement learning, and genetic algorithms.
13. The video decoding apparatus of any one of claims 9 to 12, wherein the processor is further configured to: Train the model to generate partition grouping predictions; and The partition groups are determined based on a trained model.
14. The video decoding device of any one of claims 9 to 13, wherein the partition information indicates a first partition group, a second partition group, a first partition associated with the first partition group, a second partition associated with the second partition group, first encoding technology information associated with the first partition group, second encoding technology information associated with the second partition group, a first representation parameter associated with the first partition group, and a second representation parameter associated with the second partition group.
15. The video decoding apparatus of any one of claims 9 to 14, wherein the processor is further configured to: Video data is decoded, wherein the partition information associated with the video is obtained based on the decoded video data, wherein the video data includes an indication of using a group of coding units.
16. The video decoding apparatus of any one of claims 9 to 15, wherein the processor is further configured to: The video data is decoded, wherein the partition information associated with the video is obtained based on the decoded video data, wherein the video data includes an index associated with group information.
17. A video encoding method, the video encoding method comprising: Determine a first partition associated with the input signal and a second partition associated with the input signal; Determine a first partition group and a second partition group, wherein the first partition group includes at least the first partition, and wherein the second partition group includes at least the second partition; Determine a first representation parameter associated with the first partition group and a second representation parameter associated with the second partition group; as well as The first partition, the second partition, the first partition group, the second partition group, the first representation parameter, and the second representation parameter are encoded. as well as The video data includes an encoded first partition, an encoded second partition, an encoded first partition group, an encoded second partition group, an encoded first representation parameter, and an encoded second representation parameter.
18. The video encoding method of claim 17, wherein the method further comprises: The distance between the first partition and the second partition is determined, wherein the first partition group and the second partition group are determined based on the distance between the first partition and the second partition.
19. The video coding method of any one of claims 17 and 18, wherein the method further comprises: The impact associated with partition groups is determined, wherein the impact associated with partition groups is determined based on the optimization of the coding performance associated with partition groups, wherein the first partition group and the second partition group are determined based on the impact associated with partition groups.
20. The video coding method of claim 19, wherein the optimization of coding performance associated with partitioning and grouping includes using one or more of greedy search, brute force methods, reinforcement learning, and genetic algorithms.
21. The video encoding method according to any one of claims 17 to 20, wherein the method further comprises: The model is trained to generate partition grouping predictions, wherein the first partition group and the second partition group are determined based on the trained model.
22. The video encoding method according to any one of claims 17 to 21, wherein the method further comprises: Determine a first coding technique associated with the first partition group and a second coding technique associated with the second partition group; as well as The video data includes the first encoding technique and the second encoding technique.
23. The video encoding method of any one of claims 17 to 22, wherein the video data includes an indication of using a group of encoding units.
24. The video encoding method of any one of claims 17 to 23, wherein the video data further includes a group information index, wherein the group information index indicates group information associated with the first partition and the second partition.
25. A video decoding method, the video decoding method comprising: Obtain the partition information associated with the video, wherein the partition information indicates partition group information; The decoding technique for the partition group is determined based on the partition information; The signal values associated with the partition group are determined based on the decoding technology. as well as The signal is reconstructed based on the determined signal values.
26. The video decoding method of claim 25, wherein the method further comprises: Determine the distance between the first partition and the second partition; as well as The partition group is determined based on the distance between the first partition and the second partition.
27. The video decoding method of any one of claims 25 and 26, wherein the method further comprises: Determine the impact associated with partition grouping, wherein the impact associated with partition grouping is determined based on optimization of the coding performance associated with partition grouping; as well as The partition group is determined based on the effects associated with the partition grouping.
28. The video decoding method of claim 27, wherein the optimization of the coding performance associated with partitioning and grouping includes using one or more of greedy search, brute force methods, reinforcement learning, and genetic algorithms.
29. The video decoding method according to any one of claims 25 to 28, wherein the method further comprises: Train the model to generate partition grouping predictions; as well as The partition groups are determined based on a trained model.
30. The video decoding method of any one of claims 25 to 29, wherein the partition information indicates a first partition group, a second partition group, a first partition associated with the first partition group, a second partition associated with the second partition group, first encoding technology information associated with the first partition group, second encoding technology information associated with the second partition group, a first representation parameter associated with the first partition group, and a second representation parameter associated with the second partition group.
31. The video decoding method according to any one of claims 25 to 30, wherein the method further comprises: Video data is decoded, wherein the partition information associated with the video is obtained based on the decoded video data, wherein the video data includes an indication of using a group of coding units.
32. The video decoding method according to any one of claims 25 to 31, wherein the method further comprises: The video data is decoded, wherein the partition information associated with the video is obtained based on the decoded video data, wherein the video data includes an index associated with group information.