Feature tensor compression with distribution adjustment

By adjusting and encoding/decoding eigenvalues ​​associated with feature tensors and channels, the problem of low video processing efficiency in machine vision systems is solved, achieving more efficient video encoding and decoding.

CN121844567APending Publication Date: 2026-04-10INTERDIGITAL VC HOLDINGS INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
INTERDIGITAL VC HOLDINGS INC
Filing Date
2024-09-10
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

Existing video processing technologies are inefficient for machine vision systems and cannot effectively process video and image data.

Method used

By configuring video encoding and decoding devices, and adjusting and encoding/decoding the feature values ​​associated with the channels using feature tensors, distribution alignment is achieved, including the processing of quantized and non-quantized feature values.

Benefits of technology

It improves the video processing efficiency of machine vision systems, optimizes the encoding and decoding process of feature tensors, and adapts to the needs of machine vision applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121844567A_ABST
    Figure CN121844567A_ABST
Patent Text Reader

Abstract

Video encoding and decoding devices are described herein. A video decoding device may obtain a feature tensor from a bitstream, which may include a first feature channel representing a first set of image features. The video decoding device may determine a first distribution alignment value associated with the first feature channel based on the bitstream, wherein the first distribution alignment value may represent a most frequent binary element (binary unit) associated with the first feature channel, an average value associated with the first feature channel, or a median value associated with the first feature channel. The video decoding apparatus may adjust a first feature channel of the feature tensor based on the first distribution alignment value, and perform a decoding operation using the feature tensor including the adjusted first feature channel.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-reference to related applications This application claims the benefit of U.S. Provisional Patent Application No. 63 / 537,709, filed September 11, 2023, the disclosure of which is incorporated herein by reference in its entirety. Background Technology

[0002] With the rise of machine learning techniques for vision applications, machines' consumption of video and images is likely to increase. Conventional video processing techniques (e.g., video encoding or decoding) may be suitable for human video and image consumption, but may be inefficient for machine vision-based systems or applications. Summary of the Invention

[0003] This document discloses systems, methods, and means associated with video encoding and / or decoding. The video encoding apparatus described herein can be configured to obtain a feature tensor, wherein the feature tensor may be associated with one or more channels and may represent feature values ​​associated with the one or more channels. The video encoding apparatus can also be configured to determine a value and adjust the feature tensor based on the determined value (e.g., the adjustment may include subtracting the determined value from the feature values ​​associated with the one or more channels). The video encoding apparatus can then encode the adjusted feature tensor. In an example, if the feature tensor is a quantized feature tensor, the determined value may include the most frequent feature value or binary unit (bin) associated with the one or more channels. In an example, if the feature tensor is a non-quantized feature tensor, the determined value may include the average feature value or binary unit associated with the one or more channels. In an example, as a result of the adjustment, the distribution of feature values ​​associated with the one or more channels may be zero-centered.

[0004] The video decoding apparatus described herein can be configured to obtain a feature tensor, wherein the feature tensor may be associated with one or more channels and may represent feature values ​​associated with the one or more channels. The video decoding apparatus can also be configured to determine a value. This value may be determined based on an indication in the video data (e.g., by feature tensor adjustment indication in the decoded bitstream). This value may be a quantized average, a non-quantized average, a most frequent binary unit, etc., associated with the one or more channels. The video decoding apparatus may adjust the feature tensor based on the determined value (e.g., the adjustment may include adding the determined value to the feature values ​​associated with the one or more channels). The video decoding apparatus can then decode the adjusted feature tensor. In the example, if the feature tensor is a quantized feature tensor, the determined value may include the most frequent feature value or binary unit associated with the one or more channels. In the example, if the feature tensor is a non-quantized feature tensor, the determined value may include the average feature value or binary unit associated with the one or more channels. In the example, as a result of the adjustment, the distribution of feature values ​​associated with the one or more channels may be zero-centered.

[0005] In embodiments of this disclosure, a video decoding device can be configured to: obtain a feature tensor from a bitstream, the feature tensor including a first feature channel representing a first set of image features. The video decoding device can determine a first distribution alignment value associated with the first feature channel based on the bitstream, wherein the first distribution alignment value can represent the most frequent binary symbol or binary element (binary unit) associated with the first feature channel, the average value associated with the first feature channel, or the median value associated with the first feature channel. The video decoding device can adjust the first feature channel of the feature tensor based on the first distribution alignment value, and perform a decoding operation using the feature tensor including the adjusted first feature channel.

[0006] In embodiments of this disclosure, the described feature tensor may further include a second feature channel representing a second set of image features, and the video decoding device may determine a second distribution alignment value associated with the second feature channel based on the bitstream. Such a second distribution alignment value may represent the most frequent binary unit associated with the second feature channel, the average value associated with the second feature channel, or the median value associated with the second feature channel, and the video decoding device may adjust the second feature channel of the feature tensor based on the second distribution alignment value.

[0007] In embodiments of this disclosure, the feature tensor used to perform the decoding operation may further include an adjusted second feature channel. The first and second feature channels obtained from the bitstream may have a zero-centered distribution, while the adjusted first and second feature channels may have a non-zero-centered distribution. In embodiments of this disclosure, when adjusting the first feature channel of the feature tensor based on a first distribution alignment value, the video decoding device may add the first distribution alignment value to the first feature channel obtained from the bitstream. In some examples, the video decoding device may also rescale the first feature channel of this disclosure before adjusting it based on the first distribution alignment value.

[0008] In embodiments of this disclosure, if the first feature channel is quantized in the bitstream, the first distribution alignment value may represent the most frequent binary unit associated with the first feature channel. If the first feature channel is not quantized in the bitstream, the first distribution alignment value may represent the average or median value associated with the first feature channel.

[0009] In embodiments of this disclosure, a first distribution alignment value may be encoded in a bitstream, and the feature tensor may be part of a neural network model having a split architecture between a video decoding device and an encoding device that generates the bitstream.

[0010] In embodiments of this disclosure, a video encoding apparatus can be configured to derive a feature tensor, which may include a first feature channel representing a first set of image features. The video encoding apparatus can determine a first distribution alignment value associated with the first feature channel, wherein the first distribution alignment value may represent the most frequent binary unit associated with the first feature channel, the average value associated with the first feature channel, or the median value associated with the first feature channel. The video encoding apparatus can adjust the first feature channel of the feature tensor based on the first distribution alignment value, and encode at least one of the first distribution alignment value or the feature tensor associated with the first feature channel.

[0011] In embodiments of this disclosure, the described feature tensor may further include a second feature channel representing a second set of image features, and the video encoding device may determine a second distribution alignment value associated with the second feature channel. Such a second distribution alignment value may represent the most frequent binary unit associated with the second feature channel, the average value associated with the second feature channel, or the median value associated with the second feature channel, and the video encoding device may adjust the second feature channel of the feature tensor based on the second distribution alignment value.

[0012] In embodiments of this disclosure, as a result of adjustments to the first and second feature channels, the corresponding distributions of the first and second feature channels can be aligned. For example, as a result of adjustments to the first and second feature channels, the corresponding distributions of the first and second feature channels can be centered at zero. Attached Figure Description

[0013] Figure 1A This is a system diagram illustrating an example communication system in which one or more of the disclosed embodiments can be implemented.

[0014] Figure 1B The illustration shows a method according to one embodiment. Figure 1A The diagram shows a system diagram of an example wireless transmit / receive unit (WTRU) used in a communication system.

[0015] Figure 1C The illustration shows a method according to one embodiment. Figure 1A The diagram illustrates a system diagram of an example radio access network (RAN) and an example core network (CN) used within a communication system.

[0016] Figure 1D The illustration shows a method according to one embodiment. Figure 1A The diagram shows another example RAN and another example CN used in the communication system.

[0017] Figure 2 This is a diagram illustrating an example video encoder.

[0018] Figure 3 This is a diagram illustrating an example video decoder.

[0019] Figure 4 It is a diagram illustrating an example of a system in which various aspects and examples can be implemented.

[0020] Figure 5 This is a diagram illustrating an example of a machine-oriented video coding (VCM) pipeline.

[0021] Figure 6 This is a diagram illustrating an example of a VCM pipeline.

[0022] Figure 7 This is a diagram illustrating an example of a region-based convolutional neural network (RCNN).

[0023] Figure 8 This is a diagram illustrating an example of the shape of a feature tensor.

[0024] Figure 9This is a diagram illustrating an example of a normalized overlaid histogram of independent channels.

[0025] Figure 10 This is a diagram illustrating an example of a feature tensor encoder.

[0026] Figure 11 This is a diagram illustrating an example of a superimposed histogram of quantized independent channels.

[0027] Figure 12 This is a diagram illustrating an example of a binarization module.

[0028] Figure 13 This is a diagram illustrating an example of a feature tensor decoder.

[0029] Figure 14 This is a flowchart illustrating an example of distribution alignment at the encoder.

[0030] Figure 15 This is a flowchart illustrating an example of a backshift process without dequantization at the decoder.

[0031] Figure 16 This is a flowchart illustrating an example of a dequantized shift-back process at the decoder. Detailed Implementation

[0032] A more detailed understanding can be obtained from the following description, which is given by way of example in conjunction with the accompanying drawings.

[0033] Figure 1A This is a schematic diagram illustrating an example communication system 100 in which one or more of the disclosed embodiments may be implemented. The communication system 100 may be a multi-access system that provides content such as voice, data, video, messages, and broadcasts to multiple wireless users. The communication system 100 enables multiple wireless users to access such content by sharing system resources, including wireless bandwidth. For example, the communication system 100 may employ one or more channel access methods, such as Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Frequency Division Multiple Access (FDMA), Orthogonal FDMA (OFDMA), Single Carrier FDMA (SC-FDMA), Zero Tail Unique Word DFT Spread Spectrum OFDM (ZT UWDTS-s OFDM), Unique Word OFDM (UW-OFDM), Resource Block Filtered OFDM, Filter Bank Multicarrier (FBMC), etc.

[0034] like Figure 1AAs shown, the communication system 100 may include wireless transmit / receive units (WTRUs) 102a, 102b, 102c, 102d, RAN 104 / 113, CN 106 / 115, Public Switched Telephone Network (PSTN) 108, Internet 110, and other networks 112. However, it should be understood that the disclosed embodiments contemplate any number of WTRUs, base stations, networks, and / or network elements. Each of the WTRUs 102a, 102b, 102c, and 102d may be any type of device configured to operate and / or communicate in a wireless environment. For example, WTRUs 102a, 102b, 102c, and 102d (any of which may be referred to as a “station” and / or “STA”) may be configured to transmit and / or receive wireless signals and may include user equipment (UE), mobile stations, fixed or mobile subscriber units, subscription-based units, pagers, cellular phones, personal digital assistants (PDAs), smartphones, laptops, netbooks, personal computers, wireless sensors, hotspots or Mi-Fi devices, Internet of Things (IoT) devices, watches or other wearable devices, head-mounted displays (HMDs), vehicles, drones, medical devices and applications (e.g., remote surgery), industrial devices and applications (e.g., robots and / or other wireless devices operating in industrial and / or automated processing chain environments), consumer electronics devices, devices operating on commercial and / or industrial wireless networks, etc. Any of WTRUs 102a, 102b, 102c, and 102d may be interchangeably referred to as a UE.

[0035] The communication system 100 may also include base station 114a and / or base station 114b. Each of base stations 114a and 114b may be any type of device configured to wirelessly interface with at least one of WTRUs 102a, 102b, 102c, and 102d to facilitate access to one or more communication networks, such as CN 106 / 115, the Internet 110, and / or other networks 112. For example, base stations 114a and 114b may be base transceiver stations (BTS), node B, eNode B, home node B, home eNode B, gNB, NR node B, site controller, access point (AP), wireless router, etc. Although base stations 114a and 114b are each depicted as a single element, it should be understood that base stations 114a and 114b may include any number of interconnected base stations and / or network elements.

[0036] Base station 114a may be part of RAN 104 / 113, which may also include other base stations and / or network elements (not shown), such as base station controllers (BSCs), radio network controllers (RNCs), relay nodes, etc. Base station 114a and / or base station 114b may be configured to transmit and / or receive radio signals on one or more carrier frequencies, which may be referred to as cells (not shown). These frequencies may be in licensed spectrum, unlicensed spectrum, or a combination of licensed and unlicensed spectrum. A cell may provide coverage of a specific geographic area, which may be relatively fixed or may change over time. The cell may be further divided into cell sectors. For example, the cell associated with base station 114a may be divided into three sectors. Therefore, in one embodiment, base station 114a may include three transceivers, i.e., one transceiver for each sector of the cell. In one embodiment, base station 114a may employ multiple-input multiple-output (MIMO) technology and may use multiple transceivers for each sector of the cell. For example, beamforming can be used to transmit and / or receive signals in a desired spatial direction.

[0037] Base stations 114a and 114b can communicate with one or more of WTRUs 102a, 102b, 102c, and 102d via air interface 116. Air interface 116 can be any suitable wireless communication link (e.g., radio frequency (RF), microwave, centimeter wave, micrometer wave, infrared (IR), ultraviolet (UV), visible light, etc.). Any suitable radio access technology (RAT) can be used to establish air interface 116.

[0038] More specifically, as described above, the communication system 100 can be a multi-access system and can employ one or more channel access schemes, such as CDMA, TDMA, FDMA, OFDMA, SC-FDMA, etc. For example, base stations 114a and WTRUs 102a, 102b, and 102c in RAN 104 / 113 can implement radio technologies such as Universal Mobile Telecommunications System (UMTS) Terrestrial Radio Access (UTRA), which can establish air interfaces 115 / 116 / 117 using Wideband CDMA (WCDMA). WCDMA can include communication protocols such as High-Speed ​​Packet Access (HSPA) and / or evolved HSPA (HSPA+). HSPA can include High-Speed ​​Downlink (DL) Packet Access (HSDPA) and / or High-Speed ​​UL Packet Access (HSUPA).

[0039] In one embodiment, base station 114a and WTRUs 102a, 102b, 102c may implement radio technologies such as evolved UMTS terrestrial radio access (E-UTRA), which may use Long Term Evolution (LTE) and / or Advanced LTE (LTE-A) and / or Advanced LTE Pro (LTE-A Pro) to establish air interface 116.

[0040] In one embodiment, base station 114a and WTRUs 102a, 102b, 102c can implement radio technologies such as NR radio access, which can establish an air interface 116 using a new radio (NR).

[0041] In one embodiment, base station 114a and WTRUs 102a, 102b, and 102c can implement multiple radio access technologies. For example, base station 114a and WTRUs 102a, 102b, and 102c can jointly implement LTE radio access and NR radio access, for example, using the dual connectivity (DC) principle. Therefore, the air interface used by WTRUs 102a, 102b, and 102c can be characterized by multiple types of radio access technologies and / or transmissions sent to / from multiple types of base stations (e.g., eNBs and gNBs).

[0042] In other embodiments, base station 114a and WTRUs 102a, 102b, 102c can implement radio technologies such as IEEE 802.11 (i.e., Wi-Fi), IEEE 802.16 (i.e., WiMAX), CDMA2000, CDMA2000 1X, CDMA2000 EV-DO, Provisional Standard 2000 (IS-2000), Provisional Standard 95 (IS-95), Provisional Standard 856 (IS-856), Global System for Mobile Communications (GSM), Enhanced Data Rate GSM Evolution (EDGE), GSMEDGE (GERAN), etc.

[0043] For example, Figure 1ABase station 114b can be a wireless router, home node B, home eNodeB, or access point, and can utilize any suitable RAT to facilitate wireless connectivity in a local area, such as commercial locations, homes, vehicles, campuses, industrial facilities, air corridors (e.g., for drone use), roads, etc. In one embodiment, base station 114b and WTRUs 102c, 102d can implement radio technologies such as IEEE 802.11 to establish a wireless local area network (WLAN). In one embodiment, base station 114b and WTRUs 102c, 102d can implement radio technologies such as IEEE 802.15 to establish a wireless personal area network (WPAN). In yet another embodiment, base station 114b and WTRUs 102c, 102d can utilize cellular-based RATs (e.g., WCDMA, CDMA2000, GSM, LTE, LTE-A, LTE-A Pro, NR, etc.) to establish picocells or femtocells. Figure 1A As shown, base station 114b can be directly connected to the Internet 110. Therefore, base station 114b may not need to access the Internet 110 via CN 106 / 115.

[0044] RAN 104 / 113 can communicate with CN 106 / 115, which can be any type of network configured to provide voice, data, application, and / or Voice over Internet Protocol (VoIP) services to one or more of WTRUs 102a, 102b, 102c, and 102d. Data can have different Quality of Service (QoS) requirements, such as different throughput requirements, latency requirements, fault tolerance requirements, reliability requirements, data throughput requirements, mobility requirements, etc. CN 106 / 115 can provide call control, billing services, location-based services, prepaid calling, internet connectivity, video distribution, and / or perform advanced security functions such as user authentication. Although in Figure 1A Although not shown, it should be understood that RAN104 / 113 and / or CN 106 / 115 can communicate directly or indirectly with other RANs that use the same RAT as or a different RAT than RAN 104 / 113. For example, in addition to being connected to RAN 104 / 113, which may utilize NR radio technology, CN 106 / 115 can also communicate with another RAN (not shown) that uses GSM, UMTS, CDMA 2000, WiMAX, E-UTRA, or WiFi radio technology.

[0045] CN 106 / 115 can also serve as a gateway for WTRU 102a, 102b, 102c, 102d to access PSTN 108, the Internet 110, and / or other networks 112. PSTN 108 may include a circuit-switched telephone network providing Common Old-Style Telephone Service (POTS). The Internet 110 may include a global system of interconnected computer networks and devices using common communication protocols such as Transmission Control Protocol (TCP), User Datagram Protocol (UDP), and / or Internet Protocol (IP) from the TCP / IP Internet Protocol suite. Network 112 may include wired and / or wireless communication networks owned and / or operated by other service providers. For example, network 112 may include another CN connected to one or more RANs, which may use the same RAT as RAN 104 / 113 or a different RAT.

[0046] Some or all of the WTRUs 102a, 102b, 102c, and 102d in the communication system 100 may include multi-mode capabilities (e.g., WTRUs 102a, 102b, 102c, and 102d may include multiple transceivers for communicating with different wireless networks via different wireless links). For example... Figure 1A The WTRU 102c shown can be configured to communicate with base station 114a, which may employ cellular-based radio technology, and to communicate with base station 114b, which may employ IEEE 802 radio technology.

[0047] Figure 1B This is a system diagram illustrating example WTRU 102. (Example:) Figure 1B As shown, among other things, WTRU 102 may include, in particular, a processor 118, a transceiver 120, a transmit / receive element 122, a speaker / microphone 124, a keypad 126, a display / touchpad 128, non-removable memory 130, removable memory 132, a power supply 134, a global positioning system (GPS) chipset 136, and / or other peripheral devices 138, etc. It should be understood that WTRU 102 may include any sub-combination of the foregoing elements while remaining consistent with the embodiments.

[0048] Processor 118 may be a general-purpose processor, a special-purpose processor, a conventional processor, a digital signal processor (DSP), multiple microprocessors, one or more microprocessors associated with a DSP core, a controller, a microcontroller, an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) circuit, any other type of integrated circuit (IC), a state machine, etc. Processor 118 may perform signal encoding, data processing, power control, input / output processing, and / or any other functions that enable WTRU 102 to operate in a wireless environment. Processor 118 may be coupled to transceiver 120, which may be coupled to transmitting / receiving element 122. Although Figure 1B The processor 118 and transceiver 120 are depicted as separate components, but it should be understood that the processor 118 and transceiver 120 may be integrated together in an electronic package or chip.

[0049] Transmitting / receiving element 122 can be configured to transmit signals to or receive signals from a base station (e.g., base station 114a) over air interface 116. For example, in one embodiment, transmitting / receiving element 122 can be an antenna configured to transmit and / or receive RF signals. In one embodiment, transmitting / receiving element 122 can be, for example, a transmitter / detector configured to transmit and / or receive IR, UV, or visible light signals. In yet another embodiment, transmitting / receiving element 122 can be configured to transmit and / or receive both RF and optical signals. It should be understood that transmitting / receiving element 122 can be configured to transmit and / or receive any combination of wireless signals.

[0050] Although the transmitting / receiving element 122 is in Figure 1B While depicted as a single element, WTRU 102 may include any number of transmit / receive elements 122. More specifically, WTRU 102 may employ MIMO technology. Thus, in one embodiment, WTRU 102 may include two or more transmit / receive elements 122 (e.g., multiple antennas) for transmitting and receiving wireless signals on air interface 116.

[0051] Transceiver 120 can be configured to modulate signals transmitted by transmitting / receiving element 122 and demodulate signals received by transmitting / receiving element 122. As described above, WTRU 102 can have multi-mode capability. Therefore, for example, transceiver 120 may include multiple transceivers to enable WTRU 102 to communicate via multiple RATs, such as NR and IEEE 802.11.

[0052] The processor 118 of WTRU 102 can be coupled to a speaker / microphone 124, a keypad 126, and / or a display / touchpad 128 (e.g., a liquid crystal display (LCD) unit or an organic light-emitting diode (OLED) display unit) and can receive user input data therefrom. The processor 118 can also output user data to the speaker / microphone 124, keypad 126, and / or display / touchpad 128. Furthermore, the processor 118 can access and store information from any type of suitable memory, such as non-removable memory 130 and / or removable memory 132. Non-removable memory 130 may include random access memory (RAM), read-only memory (ROM), a hard disk, or any other type of memory storage device. Removable memory 132 may include a user identification module (SIM) card, memory stick, secure digital storage (SD) card, etc. In other embodiments, the processor 118 can access and store information from memory that is not physically located on WTRU 102 (e.g., a server or home computer (not shown)).

[0053] The processor 118 can receive power from the power supply 134 and can be configured to distribute and / or control power to other components in the WTRU 102. The power supply 134 can be any suitable device that powers the WTRU 102. For example, the power supply 134 may include one or more dry cell batteries (e.g., nickel-cadmium (NiCd), nickel-zinc (NiZn), nickel metal hydride (NiMH), lithium-ion (Li-ion), etc.), solar cells, fuel cells, etc.

[0054] The processor 118 may also be coupled to a GPS chipset 136, which may be configured to provide location information (e.g., longitude and latitude) about the current location of the WTRU 102. In addition to, or instead of, information from the GPS chipset 136, the WTRU 102 may receive location information on the air interface 116 from base stations (e.g., base stations 114a, 114b) and / or determine its location based on the timing of signals received from two or more nearby base stations. It should be understood that the WTRU 102 may acquire location information using any suitable location determination method while remaining consistent with the embodiments.

[0055] The processor 118 may be further coupled to other peripheral devices 138, which may include one or more software and / or hardware modules providing additional features, functions, and / or wired or wireless connectivity. For example, peripheral devices 138 may include accelerometers, electronic compasses, satellite transceivers, digital cameras (for photos and / or videos), Universal Serial Bus (USB) ports, vibration devices, television transceivers, hands-free headsets, Bluetooth® modules, FM radio units, digital music players, media players, video game player modules, internet browsers, virtual reality and / or augmented reality (VR / AR) devices, activity trackers, etc. Peripheral devices 138 may include one or more sensors, such as gyroscopes, accelerometers, Hall effect sensors, magnetometers, orientation sensors, proximity sensors, temperature sensors, time sensors; geolocation sensors, altimeters, light sensors, touch sensors, magnetometers, barometers, attitude sensors, biosensors, and / or humidity sensors.

[0056] WTRU 102 may include a full-duplex radio for which the transmission and reception of some or all signals (e.g., signals associated with specific subframes for UL (e.g., for transmission) and downlink (e.g., for reception)) may be concurrent and / or simultaneous. The full-duplex radio may include an interference management unit to reduce and / or substantially eliminate self-interference via hardware (e.g., chokes) or via signal processing by a processor (e.g., a separate processor (not shown) or via processor 118). In one embodiment, WTRU 102 may include a half-duplex radio for which the transmission and reception of some or all signals (e.g., signals associated with specific subframes for UL (e.g., for transmission) or downlink (e.g., for reception)) may be concurrent and / or simultaneous.

[0057] Figure 1C This diagram illustrates a system diagram of RAN 104 and CN 106 according to an embodiment. As described above, RAN 104 can communicate with WTRUs 102a, 102b, and 102c via air interface 116 using E-UTRA radio technology. RAN 104 can also communicate with CN 106.

[0058] RAN 104 may include eNode-Bs 160a, 160b, and 160c; however, it should be understood that RAN 104 may include any number of eNode-Bs while remaining consistent with the embodiments. eNode-Bs 160a, 160b, and 160c may each include one or more transceivers for communicating with WTRUs 102a, 102b, and 102c on air interface 116. In one embodiment, eNode-Bs 160a, 160b, and 160c may implement MIMO technology. Therefore, for example, eNode-B 160a may use multiple antennas to transmit and / or receive radio signals from WTRU 102a.

[0059] Each of the eNode-B 160a, 160b, and 160c can be associated with a specific cell (not shown) and can be configured to handle radio resource management decisions, handover decisions, user scheduling in UL and / or DL, etc. Figure 1C As shown, eNode-B 160a, 160b, and 160c can communicate with each other on the X2 interface.

[0060] Figure 1C The CN 106 shown may include a Mobility Management Entity (MME) 162, a Serving Gateway (SGW) 164, and a Packet Data Network (PDN) Gateway (or PGW) 166. While each of the foregoing elements is described as part of CN 106, it should be understood that any of these elements may be owned and / or operated by an entity other than a CN operator.

[0061] The MME 162 can connect to each of the eNode-Bs 162a, 162b, and 162c in RAN 104 via the S1 interface and can act as a control node. For example, the MME 162 can be responsible for authenticating users of WTRUs 102a, 102b, and 102c, bearer activation / deactivation, selecting a specific serving gateway during the initial attachment of WTRUs 102a, 102b, and 102c, etc. The MME 162 can provide control plane functions for handover between RAN 104 and other RANs (not shown) employing other radio technologies such as GSM and / or WCDMA.

[0062] The SGW 164 can connect to each of the eNode Bs 160a, 160b, and 160c in RAN 104 via the S1 interface. The SGW 164 can typically route and forward user data packets to / from WTRUs 102a, 102b, and 102c. The SGW 164 can perform other functions, such as anchoring the user plane during inter-eNode B handover, triggering paging when DL data is available for WTRUs 102a, 102b, and 102c, and managing and storing the context of WTRUs 102a, 102b, and 102c.

[0063] SGW 164 can connect to PGW 166, which can provide WTRU 102a, 102b, 102c with access to packet-switched networks such as Internet 110, so as to facilitate communication between WTRU 102a, 102b, 102c and IP-enabled devices.

[0064] CN 106 can facilitate communication with other networks. For example, CN 106 can provide WTRU 102a, 102b, and 102c with access to a circuit-switched network such as PSTN 108, facilitating communication between WTRU 102a, 102b, and 102c and traditional landline communication equipment. For example, CN 106 may include, or be able to communicate with, an IP gateway (e.g., an IP Multimedia Subsystem (IMS) server) that acts as an interface between CN 106 and PSTN 108. Furthermore, CN 106 can provide WTRU 102a, 102b, and 102c with access to other networks 112, which may include other wired and / or wireless networks owned and / or operated by other service providers.

[0065] Despite WTRU in Figure 1A-1D While described as a wireless terminal, it is conceivable that in some representative embodiments, such a terminal may use (e.g., temporarily or permanently) a wired communication interface with a communication network.

[0066] In a representative embodiment, another network 112 may be a WLAN.

[0067] A WLAN in Infrastructure Basic Services Set (BSS) mode can have an access point (AP) for the BSS and one or more stations (STAs) associated with the AP. The AP can access or interface with a distributed system (DS) or another type of wired / wireless network that transmits traffic to and / or out of the BSS. Traffic originating outside the BSS destined for a STA can reach and be delivered to the STA via the AP. Traffic originating from a STA destined for an external BSS can be sent to the AP for delivery to the appropriate destination. For example, traffic between STAs within the BSS can be transmitted via the AP, where the source STA can send traffic to the AP, and the AP can deliver traffic to the destination STA. Traffic between STAs within the BSS can be considered and / or referred to as peer-to-peer traffic. Peer-to-peer traffic can be transmitted between source and destination STAs (e.g., directly between them) using Direct Link Establishment (DLS). In some representative embodiments, the DLS can use 802.11e DLS or 802.11z Tunneled DLS (TDLS). A WLAN using the Standalone BSS (IBSS) mode may not have an access point (AP), and STAs within the IBSS or using the IBSS (e.g., all STAs) can communicate directly with each other. The IBSS communication mode is sometimes referred to here as an "ad-hoc" communication mode.

[0068] When using 802.11ac infrastructure operating mode or a similar operating mode, the AP can transmit beacons on a fixed channel, such as the primary channel. The primary channel can be of a fixed width (e.g., a wide bandwidth of 20 MHz) or dynamically set via signaling. The primary channel can be the operating channel of the BSS and can be used by the STA to establish a connection with the AP. In some representative embodiments, such as in an 802.11 system, Carrier Sense Multiple Access (CSMA / CA) with collision avoidance can be implemented. For CSMA / CA, each STA, including the AP, can sense the primary channel. If a particular STA senses / detects and / or determines that the primary channel is busy, that particular STA can back off. A single STA (e.g., only one station) can transmit at any given time within a given BSS.

[0069] High-throughput (HT) STAs can communicate using a 40 MHz wide channel, for example, by combining a primary 20 MHz channel with adjacent or non-adjacent 20 MHz channels.

[0070] Very High Throughput (VHT) STAs can support channels with widths of 20 MHz, 40 MHz, 80 MHz, and / or 160 MHz. 40 MHz and / or 80 MHz channels can be formed by combining consecutive 20 MHz channels. A 160 MHz channel can be formed by combining eight consecutive 20 MHz channels, or by combining two non-consecutive 80 MHz channels, which can be referred to as an 80+80 configuration. For the 80+80 configuration, after channel coding, the data passes through a segment resolver, which splits the data into two streams. Each stream can be processed separately using Inverse Fast Fourier Transform (IFFT) and time-domain processing. These streams can be mapped onto two 80 MHz channels, and the data can be transmitted by the transmitting STA. At the receiver of the receiving STA, the operation of the 80+80 configuration can be reversed, and the combined data can be sent to the Media Access Control (MAC).

[0071] 802.11af and 802.11ah support operating modes below 1 GHz. The channel operating bandwidth and carrier in 802.11af and 802.11ah are reduced compared to those used in 802.11n and 802.11ac. 802.11af supports 5 MHz, 10 MHz, and 20 MHz bandwidths in the TV whitespace (TVWS) spectrum, while 802.11ah supports 1 MHz, 2 MHz, 4 MHz, 8 MHz, and 16 MHz bandwidths using non-TVWS spectrum. According to a representative embodiment, 802.11ah can support metering-type control / machine-type communications, such as MTC devices in macro coverage areas. MTC devices may have certain capabilities, such as limited capabilities, including support for (e.g., only) certain and / or limited bandwidths. MTC devices may include batteries with a battery life exceeding a threshold (e.g., to maintain a very long battery life).

[0072] WLAN systems that can support multiple channels and channel bandwidths, such as 802.11n, 802.11ac, 802.11af, and 802.11ah, include channels that can be designated as the primary channel. The bandwidth of the primary channel can be equal to the maximum common operating bandwidth supported by all STAs in the BSS. The bandwidth of the primary channel can be set and / or limited by the STA among all STAs operating in the BSS that supports the minimum bandwidth operating mode. In the example of 802.11ah, for STAs that support (e.g., only support) the 1 MHz mode (e.g., MTC type devices), the primary channel can be 1 MHz wide, even if the AP and other STAs in the BSS support 2 MHz, 4 MHz, 8 MHz, 16 MHz, and / or other channel bandwidth operating modes. Carrier Sense and / or Network Allocation Vector (NAV) settings may depend on the status of the primary channel. If the primary channel is busy, for example, because an STA (which only supports the 1 MHz operating mode) is transmitting to the AP, the entire available band can be considered busy, even if most of the available band remains idle and can be available.

[0073] In the United States, the available frequency band for 802.11ah is from 902 MHz to 928 MHz. In South Korea, the available frequency band is from 917.5 MHz to 923.5 MHz. In Japan, the available frequency band is from 916.5 MHz to 927.5 MHz. The total available bandwidth for 802.11ah is 6 MHz to 26 MHz, depending on the country code.

[0074] Figure 1D This diagram illustrates a system diagram of RAN 113 and CN 115 according to one embodiment. As described above, RAN 113 can communicate with WTRUs 102a, 102b, and 102c via air interface 116 using NR radio technology. RAN 113 can also communicate with CN 115.

[0075] RAN 113 may include gNBs 180a, 180b, and 180c; however, it should be understood that RAN 113 may include any number of gNBs while remaining consistent with the embodiments. gNBs 180a, 180b, and 180c may each include one or more transceivers for communicating with WTRUs 102a, 102b, and 102c on air interface 116. In one embodiment, gNBs 180a, 180b, and 180c may implement MIMO technology. For example, gNBs 180a and 180b may utilize beamforming to transmit signals to and / or receive signals from gNBs 180a, 180b, and 180c. Therefore, for example, gNB 180a may use multiple antennas to transmit radio signals to and / or receive radio signals from WTRU 102a. In one embodiment, gNBs 180a, 180b, and 180c can implement carrier aggregation technology. For example, gNB 180a can transmit multiple component carriers (not shown) to WTRU 102a. A subset of these component carriers can be on unlicensed spectrum, while the remaining component carriers can be on licensed spectrum. In one embodiment, gNBs 180a, 180b, and 180c can implement Coordinated Multipoint (CoMP) technology. For example, WTRU 102a can receive coordinated transmissions from gNBs 180a and 180b (and / or gNB 180c).

[0076] WTRUs 102a, 102b, and 102c can communicate with gNBs 180a, 180b, and 180c using transmissions associated with scalable digitization. For example, the OFDM symbol spacing and / or OFDM subcarrier spacing can differ for different transmissions, different cells, and / or different portions of the radio transmission spectrum. WTRUs 102a, 102b, and 102c can communicate with gNBs 180a, 180b, and 180c using subframes or transmission time intervals (TTIs) of various or scalable lengths (e.g., containing a variable number of OFDM symbols and / or a continuously variable absolute time).

[0077] gNBs 180a, 180b, and 180c can be configured to communicate with WTRUs 102a, 102b, and 102c in standalone and / or non-standalone configurations. In standalone configuration, WTRUs 102a, 102b, and 102c can communicate with gNBs 180a, 180b, and 180c without accessing other RANs (e.g., eNode-Bs 160a, 160b, and 160c). In standalone configuration, WTRUs 102a, 102b, and 102c can utilize one or more of gNBs 180a, 180b, and 180c as mobility anchors. In standalone configuration, WTRUs 102a, 102b, and 102c can communicate with gNBs 180a, 180b, and 180c using signals in unlicensed frequency bands. In a non-standalone configuration, WTRUs 102a, 102b, and 102c can communicate / connect with gNBs 180a, 180b, and 180c, while also communicating / connecting with another RAN such as eNode-Bs 160a, 160b, and 160c. For example, WTRUs 102a, 102b, and 102c can implement DC principles to communicate substantially simultaneously with one or more gNBs 180a, 180b, and 180c, as well as one or more eNode-Bs 160a, 160b, and 160c. In a non-standalone configuration, eNode-Bs 160a, 160b, and 160c can act as mobility anchors for WTRUs 102a, 102b, and 102c, and gNBs 180a, 180b, and 180c can provide additional coverage and / or throughput for serving WTRUs 102a, 102b, and 102c.

[0078] Each of gNBs 180a, 180b, and 180c can be associated with a specific cell (not shown) and can be configured to handle radio resource management decisions, handover decisions, user scheduling in UL and / or DL, network slicing support, dual connectivity, interoperability between NR and E-UTRA, routing user plane data to User Plane Functions (UPF) 184a and 184b, and routing control plane information to Access and Mobility Management Functions (AMF) 182a and 182b, etc. Figure 1D As shown, gNB 180a, 180b, and 180c can communicate with each other on the Xn interface.

[0079] Figure 1DThe CN 115 shown may include at least one AMF 182a, 182b, at least one UPF 184a, 184b, at least one Session Management Function (SMF) 183a, 183b, and possibly a Data Network (DN) 185a, 185b. Although each of the foregoing elements is depicted as part of the CN 115, it should be understood that any of these elements may be owned and / or operated by an entity other than a CN operator.

[0080] AMF 182a and 182b can connect to one or more gNBs 180a, 180b, and 180c in RAN 113 via the N2 interface and can act as control nodes. For example, AMF 182a and 182b can be responsible for authenticating users of WTRU 102a, 102b, and 102c, supporting network slicing (e.g., handling different PDU sessions with different requirements), selecting specific SMF 183a and 183b, managing registration areas, terminating NAS signaling, mobility management, and so on. AMF 182a and 182b can use network slicing to customize CN support for WTRU 102a, 102b, and 102c based on the service types used by WTRU 102a, 102b, and 102c. For example, different network slices can be established for different use cases, such as services relying on Ultra Reliable Low Latency Time (URLLC) access, services relying on Enhanced Massive Mobile Broadband (eMBB) access, services for Machine Type Communication (MTC) access, and / or so on. AMF 162 can provide control plane functions for handover between RAN 113 and other RANs (not shown) employing other radio technologies such as LTE, LTE-A, LTE-A Pro and / or non-3GPP access technologies such as WiFi.

[0081] SMFs 183a and 183b can connect to AMFs 182a and 182b in CN 115 via the N11 interface. SMFs 183a and 183b can also connect to UPFs 184a and 184b in CN 115 via the N4 interface. SMFs 183a and 183b can select and control UPFs 184a and 184b, and configure the routing of services through UPFs 184a and 184b. SMFs 183a and 183b can perform other functions, such as managing and allocating UE IP addresses, managing PDU sessions, controlling policy enforcement and QoS, and providing downlink data notifications. PDU session types can be IP-based, non-IP-based, Ethernet-based, etc.

[0082] UPF 184a and 184b can be connected to one or more gNBs 180a, 180b, and 180c in RAN 113 via the N3 interface. This N3 interface provides WTRU 102a, 102b, and 102c with access to packet-switched networks (such as Internet 110) to facilitate communication between WTRU 102a, 102b, 102c and IP-enabled devices. UPF 184 and 184b can perform other functions such as routing and forwarding packets, enforcing user plane policies, supporting multi-homed PDU sessions, handling user plane QoS, buffering downlink packets, and providing mobility anchoring.

[0083] CN 115 can facilitate communication with other networks. For example, CN 115 may include, or be able to communicate with, an IP gateway (e.g., an IP Multimedia Subsystem (IMS) server) that acts as an interface between CN 115 and PSTN 108. Furthermore, CN 115 can provide WTRUs 102a, 102b, and 102c with access to other networks 112, which may include other wired and / or wireless networks owned and / or operated by other service providers. In one embodiment, WTRUs 102a, 102b, and 102c may be connected to local data networks (DNs) 185a and 185b via the N3 interface to UPFs 184a and 184b and the N6 interface between UPFs 184a and 184b and DNs 185a and 185b.

[0084] Given Figure 1A-1D as well as Figure 1A-1D The corresponding descriptions herein indicate that one or more of the following functions can be performed by one or more emulation devices (not shown): WTRU 102a-d, base station 114a-b, eNode-B 160a-c, MME 162, SGW 164, PGW 166, gNB 180a-c, AMF 182a-b, UPF 184a-b, SMF183a-b, DN 185a-b, and / or any other device(s) described herein. An emulation device can be one or more devices configured to emulate one or more of the functions described herein. For example, an emulation device can be used to test other devices and / or simulate network and / or WTRU functions.

[0085] Simulation devices can be designed to perform tests on one or more other devices in laboratory and / or carrier network environments. For example, one or more simulation devices can perform one or more or all functions while being fully or partially implemented and / or deployed as part of a wired and / or wireless communication network to test other devices within the communication network. One or more simulation devices can perform one or more or all functions while being temporarily implemented / deployed as part of a wired and / or wireless communication network. Simulation devices can be directly coupled to another device for testing purposes and / or can perform tests using over-the-air wireless communication.

[0086] One or more simulation devices may perform one or more functions, including all functions, rather than being implemented / deployed as part of a wired and / or wireless communication network. For example, simulation devices may be used to test test scenarios in laboratory and / or non-deployment (e.g., testing) wired and / or wireless communication networks to implement the testing of one or more components. One or more simulation devices may be test devices. Simulation devices may transmit and / or receive data using direct RF coupling and / or wireless communication via RF circuitry (e.g., which may include one or more antennas).

[0087] This application describes various aspects, including tools, features, examples, models, methods, etc. Many of these aspects are described in detail, and are generally described in a manner that may sound restrictive, at least to illustrate the individual characteristics. However, this is for the purpose of clarity and does not limit the application or scope of those aspects. In fact, all the different aspects can be combined and interchanged to provide further aspects. Furthermore, the aspects described can also be combined and interchanged with those described in earlier submissions.

[0088] The aspects described and considered in this application can be implemented in many different forms. The accompanying drawings provided herein provide some examples, but other examples are considered. The discussion of the drawings does not limit the breadth of implementations. At least one of the aspects generally relates to video encoding and decoding, and at least one other aspect generally relates to the transmission of generated or encoded bitstreams. These and other aspects can be implemented as methods, apparatus, computer-readable media (e.g., storage media) including (e.g., having stored thereon) instructions for encoding or decoding video data according to any of the described methods, and / or computer-readable storage media having stored thereon bitstreams generated according to any of the described methods. When referred to herein, bitstream can refer to transmitted data or data that is stored, generated, and / or accessed but not transmitted (e.g., non-transient data).

[0089] In this application, the terms “reconstruction” and “decoding” are used interchangeably, as are the terms “pixel” and “sample”, and the terms “image”, “picture” and “frame” are used interchangeably.

[0090] Various methods are described herein, and each method includes one or more steps or actions for implementing the described method. Unless the correct operation of the method requires a specific order of steps or actions, the order and / or use of specific steps and / or actions may be modified or combined. Furthermore, terms such as "first," "second," etc., may be used in various examples to modify elements, components, steps, operations, etc., such as "first decoding" and "second decoding," for example. Unless specifically required, the use of such terms does not imply an ordering of operations. Therefore, in this example, the first decoding does not need to be performed before the second decoding and may occur, for example, before, during, or within a time period overlapping with the second decoding.

[0091] The various methods and other aspects described in this application can be used to modify, for example... Figure 2 and Figure 3 The video encoder 200 and decoder 300 modules shown herein are, for example, decoding modules. Furthermore, the subject matter disclosed herein can be applied to, for example, any type, format, or version of video encoding (whether described in standards or recommendations, whether pre-existing or future-developed) and any extensions to such standards and recommendations. Unless otherwise indicated or technically excluded, the aspects described in this application may be used individually or in combination.

[0092] Various numerical values ​​are used in the examples described in this application. These and other specific values ​​are used for the purpose of describing the examples, and the aspects described are not limited to these specific values.

[0093] Figure 2 This is a diagram illustrating an example video encoder. Variations of the example encoder 200 are considered, but for clarity, encoder 200 is described below, without describing all anticipated variations.

[0094] Before being encoded, the video sequence may undergo pre-coding processing 201, such as applying color transformations to the input color image (e.g., converting from RGB 4:4:4 to YCbCr 4:2:0), or performing remapping on the input image components to obtain a more compression-resistant signal distribution (e.g., using histogram equalization of one of the color components). Metadata may be associated with pre-processing and attached to the bitstream.

[0095] In encoder 200, the image is encoded by encoder elements as described below. The image to be encoded is segmented 202 and processed, for example, in units of coding units (CUs). For example, each unit is encoded using either intra-frame or inter-frame mode. When a unit is encoded in intra-frame mode, it performs intra-frame prediction 260. In inter-frame mode, motion estimation 275 and compensation 270 are performed. The encoder determines 205 which of the intra-frame or inter-frame modes to use for encoding the unit and indicates the intra-frame / inter-frame decision via, for example, a prediction mode flag. For example, the prediction residual is calculated by subtracting 210 prediction blocks from the original image block.

[0096] The predicted residual is then transformed (225) and quantized (230). The quantized transform coefficients, motion vectors, and other syntax elements are entropy-coded (245) to output a bitstream. The encoder can skip the transform and directly apply quantization to the untransformed residual signal. Alternatively, the encoder can bypass both the transform and quantization, directly encoding the residual without applying either the transform or quantization process.

[0097] The encoder decodes the coded block to provide a reference for further prediction. The quantized transform coefficients are dequantized 240 and inverse transformed 250 to decode the prediction residual. The image block is reconstructed by combining the decoded prediction residual and the prediction block 255. A loop filter 265 is applied to the reconstructed image to perform, for example, deblocking / SAO (sample adaptive offset) filtering to reduce coding artifacts. The filtered image is stored in the reference image buffer (280).

[0098] Figure 3 This is a diagram illustrating an example video decoder. In the example decoder 300, the bitstream is decoded by decoder elements, as described below. The video decoder 300 typically performs operations similar to... Figure 2 The encoding process described herein is the opposite of the decoding process. Encoder 200 typically also performs video decoding as part of the encoding of video data.

[0099] Specifically, the decoder's input includes a video bitstream, which can be generated by the video encoder 200. First, entropy decoding 330 is performed on the bitstream to obtain transform coefficients, motion vectors, and other encoded information. Image segmentation information indicates how to segment the image. Therefore, the decoder can segment the image 335 based on the decoded image segmentation information. The transform coefficients are dequantized 340 and inverse transformed 350 to decode the prediction residual. By combining the decoded prediction residual and prediction block 355, image blocks are reconstructed. Prediction blocks 370 can be obtained from intra-frame prediction 360 or motion-compensated prediction (i.e., inter-frame prediction) 375. A loop filter 365 is applied to the reconstructed image. The filtered image is stored at a reference image buffer 380.

[0100] The decoded image can undergo further post-decoding processing 385, such as inverse color transformation (e.g., from YCbCr 4:2:0 to RGB 4:4:4) or inverse remapping of the remapping process performed in pre-encoding processing 201. Post-decoding processing can utilize metadata derived in the pre-encoding process and transmitted as a signal in the bitstream. In one example, the decoded image (e.g., after applying a loop filter 365 and / or after post-decoding processing 385 (if post-decoding processing is used)) can be sent to a display device for presentation to the user.

[0101] Figure 4 This is a diagram illustrating examples of systems in which the various aspects and examples described herein can be implemented. System 400 can be embodied as a device including the various components described below and configured to perform one or more aspects of the aspects described in this document. Examples of such devices include, but are not limited to, various electronic devices such as personal computers, laptop computers, smartphones, tablet computers, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected home appliances, and servers. Elements of system 400 can be embodied individually or in combination in a single integrated circuit (IC), multiple ICs, and / or discrete components. For example, in at least one example, the processing and encoder / decoder elements of system 400 are distributed across multiple ICs and / or discrete components. In various examples, system 400 is communicatively coupled to one or more other systems or other electronic devices via, for example, a communication bus or through dedicated input and / or output ports. In various examples, system 400 is configured to implement one or more aspects of the aspects described in this document.

[0102] System 400 includes at least one processor 410 configured to execute instructions loaded therein for implementing various aspects, such as those described in this document. Processor 410 may include embedded memory, input / output interfaces, and various other circuitry as known in the art. System 400 includes at least one memory 420 (e.g., a volatile memory device and / or a non-volatile memory device). System 400 includes a storage device 440 that may include non-volatile memory and / or volatile memory, including but not limited to electrically erasable programmable read-only memory (EEPROM), read-only memory (ROM), programmable read-only memory (PROM), random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), flash memory, disk drives, and / or optical disk drives. As a non-limiting example, storage device 440 may include internal storage devices, attached storage devices (including removable and non-removable storage devices), and / or network-accessible storage devices.

[0103] System 400 includes an encoder / decoder module 430 configured to, for example, process data to provide encoded or decoded video, and the encoder / decoder module 430 may include its own processor and memory. The encoder / decoder module 430 represents one or more modules that can be included in a device to perform encoding and / or decoding functions. It is well known that a device may include one or both encoding and decoding modules. Furthermore, the encoder / decoder module 430 may be implemented as a separate element of system 400, or it may be incorporated into processor 410 as a combination of hardware and software as known to those skilled in the art.

[0104] Program code to be loaded onto processor 410 or encoder / decoder 430 to execute the various aspects described in this document may be stored in storage device 440 and subsequently loaded onto memory 420 for execution by processor 410. According to various examples, one or more of processor 410, memory 420, storage device 440, and encoder / decoder module 430 may store one or more items during the execution of the processes described in this document. Such stored items may include, but are not limited to, input video, decoded video or portions of decoded video, bitstreams, matrices, variables, and intermediate or final results from processing equations, formulas, operations, and operational logic.

[0105] In some examples, the internal memory of processor 410 and / or encoder / decoder module 430 is used to store instructions and provide working memory for processing during encoding or decoding. However, in other examples, external memory of the processing device (e.g., processor 410 or encoder / decoder module 430) is used for one or more of these functions. External memory can be memory 420 and / or storage device 440, such as volatile memory and / or non-volatile flash memory. In several examples, external non-volatile flash memory is used to store, for example, the operating system of a television. In at least one example, fast external volatile memory (such as RAM) is used as working memory for video encoding and decoding operations.

[0106] Inputs to the components of system 400 can be provided through various input devices as indicated in block 445. Such input devices include, but are not limited to: (i) a radio frequency (RF) section that receives, for example, RF signals transmitted over the air by a broadcaster; (ii) component (COMP) input terminals (or a set of COMP input terminals); (iii) universal serial bus (USB) input terminals; and / or (iv) high-definition multimedia interface (HDMI) input terminals. Figure 4 Other examples not shown include composite video.

[0107] In various examples, the input device of block 445 has associated corresponding input processing elements as known in the art. For example, the RF section may be associated with elements suitable for: (i) selecting a desired frequency (also known as selecting a signal, or limiting the signal band to a band), (ii) down-converting the selected signal, (iii) further band-limiting to a narrower band to select (e.g.,) a signal band that may be referred to as a channel in some examples, (iv) demodulating the down-converted and band-limited signal, (v) performing error correction, and / or (vi) demultiplexing to select a desired data packet stream. The RF section of various examples includes one or more elements for performing these functions, such as frequency selectors, signal selectors, band limiters, channel selectors, filters, downconverters, demodulators, error correctors, and demultiplexers. The RF section may include tuners that perform various functions among these functions, including, for example, down-converting the received signal to a lower frequency (e.g., intermediate frequency or near-baseband frequency) or down-converting it to baseband. In one set-top box example, the RF section and its associated input processing elements receive RF signals transmitted via a wired (e.g., cable) medium and perform frequency selection by filtering, down-converting, and re-filtering to the desired frequency band. Various examples rearrange the order of the components described above (and others), remove some of these components, and / or add other components that perform similar or different functions. Adding components may include inserting components between existing components, such as, for example, inserting amplifiers and analog-to-digital converters. In various examples, the RF section includes an antenna.

[0108] USB and / or HDMI terminals may include corresponding interface processors for connecting system 400 to other electronic devices across USB and / or HDMI connections. It is to be understood that various aspects of input processing (e.g., Reed-Solomon error correction) may be implemented as needed, for example, within a separate input processing IC or within processor 410. Similarly, various aspects of USB or HDMI interface processing may be implemented as needed, either within a separate interface IC or within processor 410. Demodulation, error correction, and demultiplexing streams are provided to various processing elements, including, for example, processor 410 and encoder / decoder 430, which operates in conjunction with memory and storage elements to process the data stream as needed for presentation on an output device.

[0109] Various components of system 400 can be provided within an integrated housing in which various components can be interconnected and transmit data therebetween using a suitable connection arrangement 425 (e.g., internal buses as known in the art, including inter-IC (I2C) buses, wiring and printed circuit boards).

[0110] System 400 includes a communication interface 450 that enables communication with other devices via a communication channel 460. The communication interface 450 may include, but is not limited to, a transceiver configured to transmit and receive data via the communication channel 460. The communication interface 450 may include, but is not limited to, a modem or network interface card (NIC), and the communication channel 460 may be implemented, for example, within a wired and / or wireless medium.

[0111] In various examples, wireless networks, such as Wi-Fi networks (e.g., IEEE 802.11, where IEEE stands for Institute of Electrical and Electronics Engineers), are used to stream or otherwise provide data to system 400. In these examples, the Wi-Fi signal is received via a communication channel 460 and a communication interface 450 suitable for Wi-Fi communication. The communication channel 460 in these examples is typically connected to an access point or router that provides access to external networks, including the Internet, to allow streaming applications and other over-the-top communications. Other examples use a set-top box to provide streaming data to system 400, delivering data via an HDMI connection to input block 445. Still other examples use an RF connection to input block 445 to provide streaming data to system 400. As indicated above, various examples provide data in a non-streaming manner. Furthermore, various examples use wireless networks other than Wi-Fi, such as cellular networks or Bluetooth® networks.

[0112] System 400 can provide output signals to various output devices, including display 475, speaker 485, and other peripheral devices 495. Various examples of display 475 include one or more of, for example, touchscreen displays, organic light-emitting diode (OLED) displays, curved displays, and / or foldable displays. Display 475 can be used in televisions, tablet computers, laptop computers, cellular phones (mobile phones), or other devices. Display 475 can also be integrated with other components (e.g., as in a smartphone) or separate (e.g., an external monitor for a laptop computer). In various examples, other peripheral devices 495 include one or more of stand-alone digital video discs (or digital universal discs) (DVDs, for both terms), disk players, stereo systems, and / or lighting systems. Various examples use one or more peripheral devices 495 that provide functionality based on the output of system 400. For example, a disk player performs the function of playing the output of system 400.

[0113] In various examples, signaling (such as AV.Link, Consumer Electronics Control (CEC), or other communication protocols that enable device-to-device control with or without user intervention) is used to transmit control signals between system 400 and display 475, speaker 485, or other peripheral devices 495. Output devices can be communicatively coupled to system 400 via dedicated connections through corresponding interfaces 470, 480, and 490. Alternatively, output devices can be connected to system 400 via communication interface 450 using communication channel 460. In electronic devices (such as, for example, televisions), display 475 and speaker 485 can be integrated into a single unit with other components of system 400. In various examples, display interface 470 includes display drivers, such as, for example, a timing controller (TCon) chip.

[0114] For example, if the RF section of input 445 is part of a separate set-top box, then display 475 and speaker 485 can alternatively be separated from one or more other components. In various examples where display 475 and speaker 485 are external components, output signals can be provided via dedicated output connections, including, for example, HDMI ports, USB ports, or COMP outputs.

[0115] The example can be executed by computer software implemented via processor 410, or by hardware, or by a combination of hardware and software. As a non-limiting example, the example can be implemented by one or more integrated circuits. Memory 420 can be of any type suitable for the technical environment and can be implemented using any suitable data storage technology (as a non-limiting example, such as optical storage devices, magnetic storage devices, semiconductor-based memory devices, fixed memory, and removable memory). Processor 410 can be of any type suitable for the technical environment and, as a non-limiting example, can encompass one or more of microprocessors, general-purpose computers, special-purpose computers, and processors based on multi-core architectures.

[0116] Various implementations include decoding. As used herein, “decoding” can encompass all or part of a process performed on a received encoded sequence to produce a final output suitable for display. In various examples, such a process includes one or more processes typically performed by a decoder, such as entropy decoding, inverse quantization, inverse transform, and differential decoding. In various examples, such a process also includes, or alternatively includes, processes performed by a decoder as described in the various implementations of this application.

[0117] As further examples, in one example, "decoding" refers only to entropy decoding; in another example, "decoding" refers only to differential decoding; and in yet another example, "decoding" refers to a combination of entropy decoding and differential decoding. Whether the phrase "decoding process" is intended to specifically refer to a subset of operations or generally to a broader decoding process will be clear based on the specific context of the description and is considered well understood by those skilled in the art.

[0118] Various implementations include encoding. In a manner similar to the discussion above regarding “decoding,” as used in this application, “encoding” can encompass all or part of a process performed, for example, on an input video sequence to produce an encoded bitstream. In various examples, such a process includes one or more processes typically performed by an encoder, such as segmentation, differential coding, transform, quantization, and entropy coding. In various embodiments, such a process also includes, or alternatively includes, processes performed by an encoder of the various implementations described in this application.

[0119] As further examples, in one example, "encoding" refers only to entropy encoding; in another example, "encoding" refers only to differential encoding; and in yet another example, "encoding" refers to a combination of differential and entropy encoding. Whether the phrase "encoding process" is intended to specifically refer to a subset of operations or generally to a broader encoding process will be clear based on the specific context of the description and is considered well understood by those skilled in the art.

[0120] Note that the grammatical elements used in this article are descriptive terms. Therefore, they do not preclude the use of other grammatical element names.

[0121] When the accompanying drawings are presented as flowcharts, it should be understood that block diagrams of the corresponding apparatus are also provided. Similarly, when the accompanying drawings are presented as block diagrams, it should be understood that flowcharts of the corresponding methods / processes are also provided.

[0122] The implementations and aspects described herein can be implemented, for example, in methods or processes, apparatuses, software programs, data streams, or signals. Even if discussed only in the context of a single form of implementation (e.g., discussed only as a method), the implementation of the discussed features can also be implemented in other forms (e.g., apparatuses or programs). Apparatuses can be implemented, for example, in suitable hardware, software, and firmware. Methods can be implemented, for example, in a processor, which generally refers to a processing device, including, for example, a computer, microprocessor, integrated circuit, or programmable logic device. Processors also include communication devices, such as, for example, computers, cellular phones, portable / personal digital assistants (“PDAs”), and other devices that facilitate the transfer of information between end users.

[0123] References to “an example” or “an example” or “an implementation” or “an implementation”, and their variations, mean that the specific features, structures, characteristics, etc., described in connection with the example are included in at least one example. Therefore, the phrases “in an example” or “in the example” or “in an implementation” or “in the implementation”, and any other variations, appearing throughout this application, do not necessarily refer to the same example.

[0124] Furthermore, this application may relate to "determining" fragments of various information. Determining information may include one or more of, for example, estimation information, calculation information, prediction information, or information retrieved from memory. Obtaining may include receiving, retrieving, constructing, generating, and / or determining.

[0125] Furthermore, this application may relate to “accessing” fragments of various information. Accessing information may include one or more of the following: receiving information, retrieving information (e.g., retrieving information from memory), storing information, moving information, copying information, calculating information, determining information, predicting information, or estimating information.

[0126] Furthermore, this application may relate to "receiving" fragments of various information. As with "access," receiving is intended to be a broad term. Receiving information may include one or more of, for example, accessing information or retrieving information (e.g., retrieving information from memory). Moreover, "receiving" is generally referred to in one or more ways during operations such as, for example, storing information, processing information, transmitting information, moving information, copying information, erasing information, calculating information, determining information, predicting information, or estimating information.

[0127] To be understood, for example, in the cases of “A / B,” “A and / or B,” and “at least one of A and B,” the use of any of the following “ / ,” “and / or,” and “at least one of…” is intended to cover selecting only the first listed option (A), or only the second listed option (B), or selecting both options (A and B). As a further example, in the cases of “A, B, and / or C” and “at least one of A, B, and C,” such wording is intended to cover selecting only the first listed option (A), or only the second listed option (B), or only the third listed option (C), or only the first and second listed options (A and B), or only the first and third listed options (A and C), or only the second and third listed options (B and C), or selecting all three options (A, B, and C). As will be clear to those skilled in the art and related fields, this can be extended to as many entries as possible listed.

[0128] Furthermore, among other things, as used herein, the term "signaling" also refers to instructing the corresponding decoder to do something. Thus, in the examples, the same parameters are used on both the encoder and decoder sides. Therefore, for example, the encoder can transmit (explicitly signal) a specific parameter to the decoder so that the decoder can use the same specific parameter. Conversely, if the decoder already has the specific parameter as well as other parameters, signaling can be used without transmission (implicitly signaling) to simply allow the decoder to know and select the specific parameter. Bit savings are achieved in various examples by avoiding the transmission of any actual functionality. It should be understood that signaling can be implemented in many ways. For example, in various examples, one or more syntax elements, flags, etc., are used to signal information to the corresponding decoder. Although the verb form of the term "signaling" has been mentioned above, the word "signal" can also be used as a noun in this text.

[0129] As will be apparent to those skilled in the art, implementations can generate various signals formatted to carry, for example, information that can be stored or transmitted. The information may include, for example, instructions for performing a method or data generated by one of the described implementations. For example, the signal may be formatted to carry a bitstream of the described example. Such a signal may be formatted as, for example, electromagnetic waves (e.g., using the radio frequency portion of a spectrum) or formatted as a baseband signal. Formatting may include, for example, encoding a data stream and modulating a carrier wave with the encoded data stream. The information carried by the signal may be, for example, analog or digital information. As is known, the signal can be transmitted via a variety of different wired or wireless links. The signal may be stored on a processor-readable medium, or may be accessed or received from a processor-readable medium.

[0130] Numerous examples are described herein. Features of the examples may be provided individually or in any combination across various claim classes and types. Furthermore, examples may include one or more of the features, devices, or aspects described herein, individually or in any combination across various claim classes and types. For example, features described herein may be implemented in a bitstream or signal that includes information generated as described herein. This information may allow decoding, encoding, bitstream and / or decoding of the bitstream according to any of the described embodiments. For example, features described herein may be implemented by creating and / or transmitting and / or receiving and / or decoding a bitstream or signal. For example, features described herein may be implemented by a method, process, apparatus, medium storing instructions (e.g., a computer-readable medium), medium storing data, or signal. For example, features described herein may be implemented by a TV, set-top box, cellular phone, tablet computer, or other electronic device performing decoding. The TV, set-top box, cellular phone, tablet computer, or other electronic device may display (e.g., using a monitor, screen, or other type of display) the resulting image (e.g., an image reconstructed from a residual of a video bitstream). TVs, set-top boxes, cellular phones, tablets, or other electronic devices can receive signals including encoded images and perform decoding.

[0131] This document describes examples in the context of collaborative intelligence, such as remote or shared analysis of source images or videos using deep neural networks (DNNs) capable of performing computer vision tasks such as classification, object detection, object tracking, etc. However, those skilled in the art will understand that the examples are also applicable to other contexts or use cases.

[0132] With the rise of machine learning technologies for vision applications (such as intelligent transportation, smart cities, intelligent content management, etc.), the amount of video and image content consumed by machines (e.g., compared to human video and image consumption) may increase. Vision tasks can involve significant computation and can be performed on cloud systems (e.g., rather than on devices that can be used to capture source content). Video content can be transmitted to cloud systems. Video content (e.g., source data) can be compressed to fit physical bandwidth and / or storage capacity. Conventional image and video processing methods (such as video codecs) may be suitable for human image or video consumption but may be inefficient for machine vision-based systems or applications (e.g., machine vision may be insensitive to the same artifacts produced by lossy compression).

[0133] One example of enabling efficient remote analytics could be compressing the source video using methods optimized for downstream visual tasks, rather than the human visual system. The framework used to perform such remote analytics may be referred to in this paper as Machine-Oriented Video Coding (VCM), and... Figure 5 An example of such a framework is illustrated. When used in this document, the term "video" can refer to image and / or video content. The techniques described herein can be applied to image and / or video content.

[0134] The inference process of a DNN model can be broken down into multiple (e.g., two) parts, such as neural network (NN) task part 1 and NN task part 2. Figure 6 As shown, these parts can run on different devices (e.g., NN task part 1 runs on a mobile phone or camera, and NN task part 2 runs on a network or cloud). The device that captures or stores the source video can perform NN task part 1, which can be associated with extracting features from the source video. The extracted features can be transmitted and analyzed via NN task part 2 (e.g., at a remote machine or device) (e.g., remote analysis). Splitting the DNN model can alleviate at least some of the computational burden, for example, if the device that captures or stores the source video has limited capabilities (e.g., in terms of processing, memory, and / or energy). Features can be transmitted while protecting the privacy of the original content. Intermediate data or features can be (e.g., at the split point) transmitted to a remote machine to perform the second part of model inference (e.g., NN task part 2). Different techniques can be used to compress the intermediate data. These techniques can include, for example, prediction, transformation, quantization, resampling, and / or entropy coding.

[0135] The NN task parts described in this paper (e.g., NN task part 1 and / or NN task part 2) can be pre-trained, and training can be performed without the constraints associated with generating compact intermediate features. The DNN model for computer vision can be designed and / or trained to maximize the accuracy of the final task, while the autoencoder can use a loss function that can take into account reconstruction quality and / or bit rate to transmit the encoder's output. When constructing the DNN model / architecture, one or more compression aspects can be considered (e.g., in addition to computational complexity). Figure 7 An example of a faster R-CNN model / architecture is shown. This model can include a backbone that can generate feature tensors at different spatial resolutions (e.g., P2, P3, P4, P5, P6). These feature tensors can be further processed for a final task (e.g., object detection and / or segmentation). In a split-inference system, split points can divide the neural network into multiple parts (e.g., such as...). Figure 6 The NN task parts 1 and NN task parts 2 are shown in the diagram. The set of feature tensors at the split (e.g., {P2, P3, P4, P5}) can be encoded and transmitted to the decoder to complete NN task part 2.

[0136] The feature tensor at the split can include multiple (e.g., 256) channels with various spatial resolutions (e.g., due to one or more architectural configurations of the RCNN model). Depending on the original input resolution, for example, due to scaling and / or padding operations performed according to the RCNN model / architecture, the input resolution to the model (e.g., wxh) can differ from the original input resolution (e.g., ...). w org x h org )different. Figure 8 An example shape of a feature tensor that can be encoded is shown.

[0137] The channel representations associated with the feature tensors can be adapted based on entropy encoding (e.g., via various methods and / or syntactic structures). The properties of the feature values ​​in the intermediate tensors of a visual task model can differ from those in a natural image. For example, more consideration can be given to the statistical relationships between adjacent feature values ​​than to the feature values ​​themselves. In an example visual DNN, the overall distribution of the intermediate feature tensor may follow a normal distribution, but the features within individual channels may differ. Statistical and / or structural features within individual channels can be preserved to retain the original task accuracy. Encoding intermediate feature tensors can be challenging given the varying characteristics of the channels (e.g., in terms of compression ratio).

[0138] The techniques disclosed herein can be used to align the distribution of channels while maintaining (e.g., the statistical properties of each) individual channel. For example, the properties of individual channels may be susceptible to lossy coding (e.g., quantization), and the techniques disclosed herein can preserve the statistical properties of channels while performing other coding operations.

[0139] Context-Adaptive Binary Arithmetic Coding for Deep Neural Network Compression (DeepCABAC) encodes the weighted parameters learned by the DNN for entropy coding. The shape and statistical properties of the intermediate feature tensor can more closely approximate the DNN parameters compared to coefficients derived by applying discrete sine or cosine transforms to the residuals of block predictions. The techniques disclosed in this paper are not limited to DeepCABAC or its binarization process and can be applied to other CABAC implementations and binarization processes.

[0140] When compressing and / or transmitting intermediate feature tensors in a split model, the statistical and / or structural characteristics of one (e.g., each) channel can be maintained at the decoded feature tensor. Because individual channels can exhibit different characteristics, feature tensors can be difficult to compress and can have large data volumes. Figure 10 shows a superimposed histogram of channels associated with an example feature tensor. This feature tensor can be normalized between maximum and minimum values ​​across channels. For visualization purposes, the tensor can be quantized (e.g., with 8 bits). The X-axis can represent quantized binary symbols or binary elements (binary units). One (e.g., each) histogram shown in the figure can include various means and / or standard deviations.

[0141] Using various video coding techniques, one or more prediction methods can be employed to reduce temporal or spatial redundancy. The residual data obtained by subtracting the prediction from the original signal can be transformed (e.g., transformed to the DCT or DST domain) and quantized to remove high-frequency coefficients that may be less sensitive to the human visual system. Applying compression to the feature tensor (e.g., direct application) can present challenges. For example, linear prediction may fail to model dependencies between feature channels (e.g., due to a lack of continuity / correlation on adjacent feature values), leading to inaccurate predictions and increased entropy in the residual data. As another example, removing high frequencies from the transform coefficients can save bits because some of the coefficients may fall below a Laplace distribution (e.g., fall to zero). However, such techniques can introduce degradation (e.g., regarding the accuracy of the final task) because they may homogenize the features of independent channels.

[0142] By obtaining the value of one (e.g., each) channel (e.g., the most frequent value) and subtracting that value from the corresponding feature channel, the distribution of feature channels can be aligned and / or centered (e.g., centered at zero) without transformation or quantization. The examples provided in this paper illustrate the syntax and techniques associated with distribution alignment and / or decoding to obtain a reconstructed feature tensor.

[0143] The distribution of the feature tensor can be aligned and / or centered across (e.g., all) channels. A bitstream syntax can be provided (e.g., transmitted) to enable the decoder to reconstruct the feature tensor. Figure 10 An example of a feature tensor encoder is shown, which may include one or more compression modules configured to perform the distribution alignment task described herein. For example, based on the split model / architecture described herein, the feature tensor can be encoded as... ,in and The preprocessing module can be used to normalize the input feature tensor with maximum and / or minimum values ​​and quantize it to n bits. Channel suppression can be used to reduce the number of transmission channels, for example, by grouping redundant feature channels. There are no restrictions on the modules before and / or after the distribution alignment module. For a (e.g., each) layer (e.g., ... For the feature tensor at a given location, if the feature tensor is quantized (e.g., via a preprocessing module), distribution alignment may include obtaining the most frequent binary symbol or binary element (binary unit), such as the feature binary unit value for one (e.g., each) channel. If the feature channels are not quantized, distribution alignment may include obtaining the average (e.g., arithmetic mean) of the feature channels. The feature tensor may be normalized and / or quantized (e.g., with 8 bits) before distribution alignment (e.g., without restricting normalization or quantization methods that may or may not exist in the preprocessing module). With distribution alignment, binary units (e.g., the most frequent binary unit or the average of multiple binary units) can be subtracted from the (e.g., all) feature values ​​associated with the corresponding channel (e.g., in one or more feature tensors), denoted as . Figure 11 An example of a superimposed histogram of channels within a feature tensor is shown, where the distribution can be centered at zero. For example... Figure 12 As shown, DeepCABAC allows for more efficient encoding of shift feature binary units because fewer bits can be assigned to smaller values ​​near 0. Feature values ​​or binary units (e.g., the most frequent feature values ​​or binary units) can be encoded into the bitstream, for example, to correct the original distribution at the decoder. Other encoding modules or operations (e.g., spatial downscaling, quantization, and / or entropy encoding of the resulting data) can follow the distribution modules or operations, for example, without restricting the order of these modules or operations. For example, spatial downscaling and / or quantization can be performed before distribution alignment.

[0144] Figure 13 An example of a feature tensor decoder configured to perform one or more of the tasks described herein is shown. The decoding operations performed by the decoder may include those related to... Figure 10 The operations shown are the inverse operations performed by the encoder. For example, if downscaling is performed at the encoder, the corresponding upscaling can be performed using the inverse quantized feature tensor (e.g., ) as input. The upscaled feature tensor (e.g., The output of the distributed shift module can be fed into the post-processing module, allowing a set of the most frequent eigenvalues ​​after parsing to be added back to the corresponding feature channels. If there is still processing to be performed in the post-processing module, the output of the distributed shift module can be further updated, and a reconstructed feature tensor can be created as the output of the decoder.

[0145] As described in this paper, distribution alignment can be achieved by subtracting binary units (e.g., the most frequent binary units) from the quantized eigenvalues ​​to move the histogram of the zero-centered quantized feature channels. If the input feature tensor is not quantized, the computed average can be subtracted from the eigenvalues.

[0146] Figure 14 The illustration shows an example of distribution alignment at the encoder. If the number of layers is greater than 1 (e.g., {P2, P3, P4, P5}), it can be repeated on multiple (e.g., all) feature tensor layers. Figure 14 The process is shown in the diagram. For those in the layer... One (e.g., each) channel (For example, until) If the input feature tensor is quantized to n bits beforehand, the most frequent binary unit can be obtained. Otherwise, the calculations related to obtaining the most frequent binary units can be replaced by calculating the average value of the characteristic channels. It can be subtracted from the feature channel and moved to subsequent channels. If the most frequent binary unit is obtained and the subtraction on the channel is completed, then the binary unit... It can be encoded into a bitstream. In the example, It can be further quantized to reduce the number of bits associated with the encoding of the binary unit. For example, if quantization is not performed, it can be encoded using fixed-length symbols or arithmetic coding (e.g., DeepCABAC) (e.g., directly). Various prediction methods can be used to encode. (For example, to reduce the number of bits to be compressed). This can be done... The median value is sent using a signal, and it can be used to... The difference between the median value and (e.g., all) binary units is encoded. On the decoder side, the parsed median value can be added to the parsed difference to reconstruct the... .

[0147] In the example, It can be quantized without sending quantization information to the decoder via a signal. Figure 15 The diagram illustrates an example decoding process, where the parsed data can be decoded... Add the feature tensor to the decoded histogram and move it back and forth.

[0148] In the example, It can be quantized, and a flag (e.g., quantized_dc_flag) can be sent to the decoder along with the corresponding dequantization information. Figure 16 This diagram illustrates an example of the decoding process, where information about... A flag indicating whether it has been quantized. If this flag indicates the parsed result... If it is quantized, then further information about dequantization can be analyzed to... Dequantization. Dequantization It can be added to the feature tensor of the decoder. If the parsed... If it is not quantized, then the parsed... It can be (e.g., directly) added to the feature tensor of the decoder.

[0149] The techniques disclosed herein can be used with any codec or any data involving multidimensional (e.g., three-way) arrays or tensors.

[0150] Although the features and elements have been described above in specific combinations, those skilled in the art will understand that each feature or element can be used alone or in any combination with other features and elements. Furthermore, the methods described herein can be implemented in a computer program, software, or firmware incorporated in a computer-readable medium for execution by a computer or processor. Examples of computer-readable media include electronic signals (transmitted via a wired or wireless connection) and computer-readable storage media. Examples of computer-readable storage media include, but are not limited to, read-only memory (ROM), random access memory (RAM), registers, cache memory, semiconductor memory devices, magnetic media (such as internal hard disks and removable disks), magneto-optical media, and optical media (such as CD-ROMs and digital multifunction discs (DVDs)). The processor associated with the software can be used to implement a radio frequency transceiver used in a WTRU, UE, terminal, base station, RNC, or any host computer.

Claims

1. A video decoding device, comprising: The processor is configured as follows: A feature tensor is obtained from the bitstream, the feature tensor including a first feature channel representing a first set of image features; Based on the bitstream, a first distribution alignment value associated with the first feature channel is determined, wherein the first distribution alignment value represents the most frequent binary element (binary unit) associated with the first feature channel, the average value associated with the first feature channel, or the median value associated with the first feature channel. Adjust the first feature channel of the feature tensor based on the first distribution alignment value; and The decoding operation is performed using the feature tensor that includes the adjusted first feature channel.

2. The video decoding device according to claim 1, wherein the feature tensor further includes a second feature channel representing a second set of image features, and wherein the processor is further configured to: Based on the bitstream, a second distribution alignment value is determined associated with the second feature channel, wherein the second distribution alignment value represents the most frequent binary unit associated with the second feature channel, the average value associated with the second feature channel, or the median value associated with the second feature channel; and The second feature channel of the feature tensor is adjusted based on the second distribution alignment value.

3. The video decoding device according to claim 2, wherein the feature tensor for performing the decoding operation further includes an adjusted second feature channel.

4. The video decoding device according to claim 2, wherein the first feature channel and the second feature channel obtained from the bitstream have a zero-centered distribution, while the adjusted first feature channel and the adjusted second feature channel have a non-zero-centered distribution.

5. The video decoding apparatus according to any one of claims 1-4, wherein the processor is configured to adjust the first feature channel of the feature tensor based on the first distribution alignment value, comprising: The processor is configured to add the first distribution alignment value to the first feature channel obtained from the bitstream.

6. The video decoding device according to any one of claims 1-5, wherein, Before adjusting the first feature channel based on the first distribution alignment value, the processor is also configured to rescale the first feature channel.

7. The video decoding device according to any one of claims 1-6, wherein, Given that the first feature channel is quantized in the bitstream, the first distribution alignment value represents the most frequent binary unit associated with the first feature channel.

8. The video decoding device according to any one of claims 1-7, wherein, When the first feature channel is not quantized in the bitstream, the first distribution alignment value represents the average or median value associated with the first feature channel.

9. The video decoding apparatus according to any one of claims 1-8, wherein the first distribution alignment value is encoded in the bitstream.

10. The video decoding device according to any one of claims 1-9, wherein the feature tensor is part of a neural network model having an architecture split between the video decoding device and the encoding device that generates the bitstream.

11. A video decoding method, comprising: A feature tensor is obtained from the bitstream, the feature tensor including a first feature channel representing a first set of image features; Based on the bitstream, a first distribution alignment value associated with the first feature channel is determined, wherein the first distribution alignment value represents the most frequent binary element (binary unit) associated with the first feature channel, the average value associated with the first feature channel, or the median value associated with the first feature channel. Adjust the first feature channel of the feature tensor based on the first distribution alignment value; and The decoding operation is performed using the feature tensor that includes the adjusted first feature channel.

12. The video decoding method according to claim 11, wherein the feature tensor further includes a second feature channel representing a second set of image features, and wherein the video decoding method further includes: Based on the bitstream, a second distribution alignment value is determined associated with the second feature channel, wherein the second distribution alignment value represents the most frequent binary unit associated with the second feature channel, the average value associated with the second feature channel, or the median value associated with the second feature channel; and The second feature channel of the feature tensor is adjusted based on the second distribution alignment value.

13. The video decoding method according to claim 12, wherein the feature tensor for performing the decoding operation further includes an adjusted second feature channel.

14. The video decoding method of claim 12, wherein the first feature channel and the second feature channel obtained from the bitstream have a zero-centered distribution, while the adjusted first feature channel and the adjusted second feature channel have a non-zero-centered distribution.

15. The video decoding method according to any one of claims 11-14, wherein adjusting the first feature channel of the feature tensor based on the first distribution alignment value comprises: Add the first distribution alignment value to the first feature channel obtained from the bitstream.

16. The video decoding method according to any one of claims 11-15, wherein, The first feature channel is rescaled before being adjusted based on the first distribution alignment value.

17. The video decoding method according to any one of claims 11-16, wherein, Given that the first feature channel is quantized in the bitstream, the first distribution alignment value represents the most frequent binary unit associated with the first feature channel.

18. The video decoding method according to any one of claims 11-17, wherein, When the first feature channel is not quantized in the bitstream, the first distribution alignment value represents the average or median value associated with the first feature channel.

19. The video decoding method according to any one of claims 11-18, wherein the first distribution alignment value is encoded in the bitstream.

20. The video decoding method according to any one of claims 11-19, wherein the feature tensor is part of a neural network model having an architecture split between the video decoding device and the encoding device that generates the bitstream.

21. A video encoding device, comprising: The processor is configured as follows: Derive a feature tensor, the feature tensor including a first feature channel representing a first set of image features; Determine a first distribution alignment value associated with the first feature channel, wherein the first distribution alignment value represents the most frequent binary element (binary unit) associated with the first feature channel, the average value associated with the first feature channel, or the median value associated with the first feature channel; Adjust the first feature channel of the feature tensor based on the first distribution alignment value; and Encode at least one of the first distribution alignment value or the feature tensor associated with the first feature channel.

22. The video encoding apparatus of claim 21, wherein the feature tensor further includes a second feature channel representing a second set of image features, and wherein the processor is further configured to: Determine a second distribution alignment value associated with the second feature channel, wherein the second distribution alignment value represents the most frequent binary unit associated with the second feature channel, the average value associated with the second feature channel, or the median value associated with the second feature channel; and The second feature channel of the feature tensor is adjusted based on the second distribution alignment value.

23. The video encoding device according to claim 22, wherein, As a result of the adjustments made to the first and second feature channels, the corresponding distributions of the first and second feature channels are aligned.

24. The video encoding device according to claim 23, wherein, As a result of the adjustments made to the first and second feature channels, the corresponding distributions of the first and second feature channels are centered at zero.

25. A video coding method, comprising: Derive a feature tensor, the feature tensor including a first feature channel representing a first set of image features; Determine a first distribution alignment value associated with the first feature channel, wherein the first distribution alignment value represents the most frequent binary element (binary unit) associated with the first feature channel, the average value associated with the first feature channel, or the median value associated with the first feature channel; Adjust the first feature channel of the feature tensor based on the first distribution alignment value; and Encode at least one of the first distribution alignment value or the feature tensor associated with the first feature channel.

26. The video coding method of claim 25, wherein the feature tensor further includes a second feature channel representing a second set of image features, and wherein the video coding method further includes: Determine a second distribution alignment value associated with the second feature channel, wherein the second distribution alignment value represents the most frequent binary unit associated with the second feature channel, the average value associated with the second feature channel, or the median value associated with the second feature channel; and The second feature channel of the feature tensor is adjusted based on the second distribution alignment value.

27. The video encoding method according to claim 26, wherein, As a result of the adjustments made to the first and second feature channels, the corresponding distributions of the first and second feature channels are aligned.

28. The video encoding method according to claim 27, wherein, As a result of the adjustments made to the first and second feature channels, the corresponding distributions of the first and second feature channels are centered at zero.