Differential signaling for machine-oriented video coding
By using multiple neural networks and dynamically adjusting their structure and parameters during video encoding and decoding, the problem of poor bit rate in video transmission was solved, achieving more efficient video data transmission.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- QUALCOMM INC
- Filing Date
- 2024-10-03
- Publication Date
- 2026-04-24
AI Technical Summary
Existing video compression technologies struggle to adapt effectively to changes in network conditions during transmission, resulting in poor bitrates and impacting the efficiency of video data transmission.
Various types of neural networks are used for video encoding and decoding. By dynamically adjusting the structure and parameters of the neural networks, the encoding method of video data is dynamically adjusted according to network conditions, and bit rate adaptation is performed using artificial intelligence and machine learning technologies.
By dynamically adjusting the neural network, the bit rate of video data transmission is reduced, improving the transmission efficiency and quality of video data and adapting to changes in different network conditions.
Smart Images

Figure CN121925844A_ABST
Abstract
Description
[0001] This application claims priority to U.S. Patent Application No. 18 / 904,887, filed October 2, 2024, and U.S. Provisional Application No. 63 / 587,917, filed October 4, 2023, the entire contents of each of which are incorporated herein by reference. U.S. Patent Application No. 18 / 904,887, filed October 2, 2024, claims the benefit of U.S. Provisional Application No. 63 / 587,917, filed October 4, 2023. Technical Field
[0002] This disclosure relates to the transmission of encoded video data. Background Technology
[0003] Digital video capabilities can be incorporated into a wide range of devices, including digital televisions, digital direct broadcasting systems, wireless broadcasting systems, personal digital assistants (PDAs), laptops or desktop computers, digital cameras, digital recording devices, digital media players, video game devices, video game consoles, cellular or satellite broadcast phones, video conferencing equipment, and more. Digital video devices implement video compression technologies (such as those described in standards defined by MPEG-2, MPEG-4, ITU-T H.263 or ITU-T H.264 / MPEG-4, Part 10, Advanced Video Decoding (AVC), ITU-T H.265 (also known as High Efficiency Video Decoding (HEVC)), and extensions to these standards) to transmit and receive digital video information more efficiently.
[0004] Video compression techniques perform spatial and / or temporal prediction to reduce or remove inherent redundancy in video sequences. For block-based video decoding, video frames or slices can be divided into macroblocks. Each macroblock can be further subdivided. Macroblocks in intra-frame decoding (I) frames or slices can be encoded using spatial prediction of adjacent macroblocks. Macroblocks in inter-frame decoding (P or B) frames or slices can be encoded using spatial prediction of adjacent macroblocks in the same frame or slice, or temporal prediction of other reference frames.
[0005] After video data is encoded, it can be packaged for transmission or storage. Video data can be assembled into video files that conform to any of various standards, such as the International Organization for Standardization (ISO) Basic Media File Format and its extensions, such as AVC. Summary of the Invention
[0006] Generally, this disclosure describes techniques for performing bit rate adaptation when streaming video data to a task network via a network, such as a 5G radio access network (RAN). The video data may be decoded using machine-oriented video decoding (VCM) techniques. Artificial intelligence / machine learning (AI / ML) techniques, such as neural networks, may be used for encoding and decoding the video data. The video encoder may specify changes to the neural network to reconfigure the decoder-side neural network in the video decoder. The data may be transmitted in the same frequency band as the video data, for example, in a high-level syntax (HLS), such as in a parameter set, or in a supplementary enhancement information (SEI) message or a generic supplementary enhancement information (VSEI) message. Alternatively, the data may be transmitted outward to a video decoding layer, for example, in an RTP extension header of a Real-time Transport Protocol (RTP) packet carrying the video data. In this way, the video decoder can determine how to configure the decoder-side neural network to appropriately decode the video data. Therefore, various neural networks can be used, each of which can be adapted to different aspects of the video data. That is, a particular neural network can be determined to compress a specific portion of video data (e.g., region of interest) or at a specific stage of the video data (e.g., temporal resampling, spatial resampling, or post-filtering) better than other neural networks, and thus these techniques can reduce the bit rate of the bit stream used to carry the decoding video data.
[0007] In one example, a method for processing video data includes: receiving data representing a plurality of neural networks associated with a video bitstream, each of the plurality of neural networks having a different type; receiving data representing an update to at least one of the neural networks, the data including a type corresponding to the at least one neural network and a neural network structure for updating; updating the neural networks according to the data representing the update to generate an updated neural network; and providing video data from the video bitstream to the updated neural network so that the updated neural network processes the video data.
[0008] In another example, an apparatus for processing video data includes: a memory configured to store video data; and a processing system implemented in a circuit, the processing system being configured to: receive data representing a plurality of neural networks associated with a video bitstream, each of the plurality of neural networks having a different type; receive data representing an update to at least one of the neural networks, the data including a type corresponding to the at least one of the neural networks and a neural network structure for updating; update the neural networks according to the data representing the update to generate an updated neural network; and provide video data from the video bitstream to the updated neural network to enable the updated neural network to process the video data.
[0009] In another example, an apparatus for processing video data includes: components for receiving data representing a plurality of neural networks associated with a video bitstream, each of the plurality of neural networks having a different type; components for receiving data representing an update to at least one of the neural networks, the data including a type corresponding to the at least one of the neural networks and a neural network structure for updating; components for updating the neural networks according to the data representing the update to generate an updated neural network; and components for providing video data from the video bitstream to the updated neural network so that the updated neural network processes the video data.
[0010] In another example, a computer-readable storage medium stores instructions that, when executed, cause a processor to: receive data representing a plurality of neural networks associated with a video bitstream, each of the plurality of neural networks having a different type; receive data representing an update to at least one of the neural networks, the data including a type corresponding to the at least one of the neural networks and a neural network structure for the update; update the neural networks according to the data representing the update to generate an updated neural network; and provide video data from the video bitstream to the updated neural network to cause the updated neural network to process the video data.
[0011] Details of one or more examples are set forth in the accompanying drawings and the following description. Other features, objects, and advantages will be apparent from these descriptions and drawings, and from the claims. Attached Figure Description
[0012] Figure 1 This is a block diagram illustrating an example system for implementing a technology for streaming media data over a network.
[0013] Figure 2 This is a conceptual diagram illustrating an example system that can implement the technology of this disclosure.
[0014] Figure 3 This is a diagram illustrating an example calculation diagram according to the technology of this disclosure.
[0015] Figure 4 This is a flowchart illustrating an example method for updating a neural network according to the technology of this disclosure. Detailed Implementation
[0016] Neural networks can be used to encode and decode media data such as image or video data. Various neural networks can be designed to perform encoding and decoding tasks. For example, neural networks can be designed to perform region-of-interest (ROI) based decoding, neural network-based intra-frame decoding, frame-level spatial resampling, temporal resampling (for video data), and / or post-filtering. Frame-level spatial resampling can include spatially downsampling frames to obtain a lower spatial resolution on the encoder side, and then upsampling the frames on the decoder side. Temporal resampling can include dropping a specific number of frames at the encoder side (e.g., dropping three frames out of every four), and then upsampling the video data on the decoder side.
[0017] The techniques disclosed herein typically involve selecting different neural networks and / or updating neural networks during encoding and decoding. The encoder may signal to the decoder data indicating an initial set of one or more neural networks, and then subsequently transmit updates to the decoder representing different neural networks or different configurations or values of the neural networks. For example, the structure of the neural network may be changed or discarded, different sets of networks may be used, and / or different weights or bias values may be used.
[0018] Therefore, the encoder can test various neural networks and / or neural network configurations to determine the set of neural networks that result in the optimal encoding of image or video data. The decoder can then be reconfigured to appropriately decode the resulting image or video data. In this way, these techniques can reduce the bit rate of the associated bitstream containing the image or video data.
[0019] These technologies are typically based on machine learning methods, such as artificial intelligence / machine learning (AI / ML), including neural networks. These technologies can be used in conjunction with hybrid techniques, which may include both traditional engineering decoding techniques and neural network-based techniques.
[0020] Figure 1 This is a block diagram illustrating an example system 10 for implementing techniques for streaming media data over a network. In this example, system 10 includes a content preparation device 20, a server device 60, and a client device 40. Client device 40 and server device 60 are communicatively coupled via a network 74, which may include the Internet. In some examples, content preparation device 20 and server device 60 may also be coupled via network 74 or another network, or may be directly communicatively coupled. In some examples, content preparation device 20 and server device 60 may include the same device.
[0021] exist Figure 1In the example, content preparation device 20 includes an audio source 22 and a video source 24. Audio source 22 may include, for example, a microphone that generates electrical signals representing captured audio data to be encoded by audio encoder 26. Alternatively, audio source 22 may include: a storage medium storing previously recorded audio data; an audio data generator, such as a computerized synthesizer; or any other audio data source. Video source 24 may include: a video camera that generates video data to be encoded by video encoder 28; a storage medium encoding previously recorded video data; a video data generation unit, such as a computer graphics source; or any other video data source. Content preparation device 20 is not necessarily communicatively coupled to server device 60 in all examples, but multimedia content may be stored on a separate medium that is read by server device 60.
[0022] The raw audio and video data may include analog or digital data. Analog data may be digitized before being encoded by audio encoder 26 and / or video encoder 28. While a speaker is speaking, audio source 22 may obtain audio data from that speaker, and video source 24 may simultaneously obtain video data of that speaker. In other examples, audio source 22 may include a computer-readable storage medium containing stored audio data, and video source 24 may include a computer-readable storage medium containing stored video data. Thus, the techniques described in this disclosure can be applied to live, streaming, real-time audio and video data, or to archived, pre-recorded audio and video data.
[0023] An audio frame corresponding to a video frame is typically an audio frame containing audio data, which is simultaneously captured (or generated) by audio source 22 and video data captured (or generated) by video source 24 and contained within the video frame. For example, when a speaker typically generates audio data by speaking, audio source 22 captures the audio data, and video source 24 simultaneously (i.e., while audio source 22 is capturing audio data) captures the speaker's video data. Therefore, an audio frame can temporally correspond to one or more specific video frames. Thus, an audio frame corresponding to a video frame typically corresponds to the following situation: in which audio data and video data are captured simultaneously, and in this situation, the audio frame and video frame respectively include the simultaneously captured audio data and video data.
[0024] In some examples, audio encoder 26 may encode a timestamp into each encoded audio frame, where the timestamp indicates the time when the audio data for the encoded audio frame was recorded, and similarly, video encoder 28 may encode a timestamp into each encoded video frame, where the timestamp indicates the time when the video data for the encoded video frame was recorded. In such examples, the audio frame corresponding to the video frame may include: an audio frame including a timestamp, and a video frame including the same timestamp. Content preparation device 20 may include an internal clock, which audio encoder 26 and / or video encoder 28 may use to generate timestamps, or audio source 22 and video source 24 may use the internal clock to associate audio and video data with timestamps, respectively.
[0025] In some examples, audio source 22 may transmit data corresponding to the time when the audio data was recorded to audio encoder 26, and video source 24 may transmit data corresponding to the time when the video data was recorded to video encoder 28. In some examples, audio encoder 26 may encode sequence identifiers into the encoded audio data to indicate the relative time order of the encoded audio data, rather than the absolute time when the audio data was recorded; similarly, video encoder 28 may use sequence identifiers to indicate the relative time order of the encoded video data. Similarly, in some examples, sequence identifiers may be mapped or otherwise associated with timestamps.
[0026] Audio encoder 26 typically produces encoded audio data streams, while video encoder 28 produces encoded video data streams. Each individual data stream (whether audio or video) can be referred to as an elementary stream. An elementary stream is a single, digitally decoded (possibly compressed) component of a media presentation. For example, the decoded video or audio portion of a media presentation can be an elementary stream. Elementary streams can be converted into Packed Elementary Streams (PES) before being encapsulated into a video file. Within the same media presentation, stream IDs can be used to distinguish PES packets belonging to one elementary stream from those belonging to another. The basic unit of data in an elementary stream is the packed elementary stream (PES) packet. Therefore, decoded video data typically corresponds to an elementary video stream. Similarly, audio data corresponds to one or more corresponding elementary streams.
[0027] exist Figure 1In one example, the encapsulation unit 30 of the content preparation device 20 receives a base stream of video data including decoded data from the video encoder 28 and a base stream of audio data including decoded data from the audio encoder 26. In some examples, both the video encoder 28 and the audio encoder 26 may include packers for forming PES packets based on the encoded data. In other examples, both the video encoder 28 and the audio encoder 26 may interface with corresponding packers for forming PES packets based on the encoded data. In other examples, the encapsulation unit 30 may include packers for forming PES packets based on the encoded audio and video data.
[0028] Video encoder 28 is capable of encoding video data of multimedia content in various ways to produce different representations of the multimedia content at various bit rates and utilizing various characteristics such as pixel resolution, frame rate, compliance with various decoding standards, compliance with various profiles and / or profile levels used for various decoding standards, representations with one or more views (e.g., for two-dimensional or three-dimensional playback), or other such characteristics. As used in this disclosure, the representation may include one of the following: audio data, video data, text data (e.g., for closed captions), or other such data. The representation may include a primary stream, such as an audio primary stream or a video primary stream. Each PES packet may include a stream_id, which identifies the primary stream to which the PES packet belongs. Encapsulation unit 30 is responsible for assembling the primary streams into streamable media data.
[0029] According to the technology disclosed herein, video encoder 28 may include a variety of different neural networks. Additionally or alternatively, video encoder 28 may be configured to dynamically reconfigure the neural networks during an encoding task. Various neural networks may be configured to perform different encoding tasks. Therefore, video encoder 28 may include a region-of-interest (ROI) based neural network encoder, an intra-frame prediction neural network encoder, an inter-frame prediction neural network encoder, a frame-level spatial resampling neural network, a temporal resampling neural network, and / or a post-filtering neural network. Similarly, each neural network may include various configurable parameters, such as weights and biases. In some examples, alternative neural networks with different structures (e.g., neuron configurations) or combinations / sequences of structures may exist.
[0030] The video encoder 28 can encode data representing one or more neural networks used in decoding the encoded video bitstream for the encoded video bitstream. The data representing the one or more neural networks can include data indicating the type or purpose of the neural network for each neural network (e.g., ROI-based decoding, intra-frame prediction decoding, frame-level spatial resampling, temporal resampling, or post-filtering). For each neural network, the data can further specify one or more structures of the neural network, as well as weights and biases. The data can also specify an identifier for each neural network, such as a sequence number.
[0031] The video encoder 28 may further encode data representing updates to one or more neural networks in the neural network. Updates may include the purpose or type of the neural network to be updated, the identifier of the neural network (e.g., an update sequence number corresponding to the current update), changes to the structure of the neural network, changes to one or more weights of the neural network, changes to one or more bias values of the neural network, new values for configuration parameters, or changes to configuration parameters (e.g., adding or removing configuration parameters). Configuration parameters may include, for example, quantization parameters, spatial downsampling ratios, and / or temporal downsampling ratios.
[0032] The video encoder 28 may signal data representing the neural network and / or updates to the neural network in High-Level Syntax (HLS) video data (such as in Sequence Parameter Set (SPS), Picture Parameter Set (PPS), or Adaptation Parameter Set (APS)) or in Supplemental Enhancement Layer (SEI) messages. Additionally or alternatively, the output interface 32 may be configured to specify data representing the neural network and / or updates to the neural network in network data external to the HLS data and video data themselves. That is, the HLS data and video data can be considered application layer data in the OSI network model, and the output interface 32 may specify data representing the neural network and / or updates to the neural network in program data units of application protocols or transport protocols (e.g., Real-Time Transport Protocol (RTP) packets), in the payload of RTP packets separate from the video data itself, in HTTP packets, etc. Alternatively, the output interface 32 may specify data representing the neural network and / or updates to the neural network in the RTP extension header of RTP packets that include the corresponding video data.
[0033] Encapsulation unit 30 receives PES packets from audio encoder 26 and video encoder 28 for media presentation of the basic stream, and forms corresponding Network Abstraction Layer (NAL) units based on the PES packets. Decoded video segments can be organized into NAL units that provide a “network-friendly” video representation for addressing applications such as video telephony, storage, broadcasting, or streaming. NAL units can be classified as Video Decoding Layer (VCL) NAL units and non-VCL NAL units. VCL units may contain the core compression engine and may include block, macroblock, and / or slice-level data. Other NAL units may be non-VCL NAL units. In some examples, a picture of the decoded data in a time instance (typically presented as a picture of the main decoded data) may be included in an access unit, which may include one or more NAL units.
[0034] Non-VCL NAL units can include parameter set NAL units and SEI NAL units, etc. Parameter sets can contain sequence-level header information (in the Sequence Parameter Set (SPS)) and picture-level header information that changes very little (in the Picture Parameter Set (PPS)). Using parameter sets (e.g., PPS and SPS), it is not necessary to repeat information that changes very little for each sequence or picture; therefore, decoding efficiency can be improved. Furthermore, the use of parameter sets allows for out-of-band transmission of important header information, thus avoiding redundant transmissions required for error recovery. In an example of out-of-band transmission, parameter set NAL units can be transmitted on a different channel than other NAL units (such as SEI NAL units).
[0035] Supplemental Enhancement Information (SEI) can contain information that is not essential for decoding image samples from VCL NAL units but can assist in processes related to decoding, display, error recovery, and other purposes. SEI messages can be included in non-VCL NAL units. SEI messages are a specification part of some standards and are therefore not always mandatory for specific implementations of standards-compliant decoders. SEI messages can be sequence-level or image-level. Some sequence-level information can be included in SEI messages, such as the scalability information SEI message in the SVC example and the view scalability information SEI message in MVC. These example SEI messages can convey information such as the extraction and characteristics of operation points.
[0036] Server device 60 includes a sending unit 70 and a network interface 72. In some examples, the sending unit 70 may be based on Real-Time Transport Protocol (RTP) and UDP. In some examples, the sending unit 70 may be based on HTTP and TCP. In some examples, the sending unit 70 may be based on HTTP, QUIC, and UDP. In some examples, server device 60 may include multiple network interfaces. Furthermore, any or all of the features of server device 60 may be implemented on other devices in the content delivery network, such as routers, bridges, proxy devices, switches, or other devices. In some examples, intermediate devices in the content delivery network may cache data of multimedia content 64 and include components substantially consistent with those of server device 60. Generally, network interface 72 is configured to transmit and receive data via network 74.
[0037] Sending unit 70 is configured to deliver media data to client device 40 via network 74 according to a network communication protocol, such as Real-Time Transport Protocol (RTP), which is standardized in Request for Comments (RFC) 3550 of the Internet Engineering Task Force (IETF). Sending unit 70 may also implement RTP-related protocols, such as RTP Control Protocol (RTCP), Real-Time Streaming Protocol (RTSP), Session Initiation Protocol (SIP), and / or Session Description Protocol (SDP). Sending unit 70 may transmit media data via network interface 72, which may implement Uniform Datagram Protocol (UDP) and / or Internet Protocol (IP). Therefore, in some examples, server device 60 may use network 74 to transmit media data via RTP and RTSP over UDP.
[0038] The sending unit 70 may receive an RTSP description request from, for example, the client device 40. The RTSP description request may include data indicating what types of data the client device 40 supports. The sending unit 70 may respond to the client device 40 with data indicating a media stream (such as media content 64), which may be transmitted to the client device 40 together with a corresponding network location identifier (such as a Uniform Resource Locator (URL) or Uniform Resource Name (URN)).
[0039] Then, sending unit 70 can receive an RTSP establishment request from client device 40. The RTSP establishment request typically indicates how the media stream will be transmitted. The RTSP establishment request may include a network location identifier and transmission specifier for the requested media data (e.g., media content 64), such as the local port used to receive RTP data and control data (e.g., RTCP data) on client device 40. Sending unit 70 can respond to the RTSP establishment request with acknowledgments and data representing the port of server device 60 through which the RTP data and control data will be transmitted. Sending unit 70 can then receive an RTSP playback request to allow the media stream to be "played," i.e., transmitted to client device 40 via network 74. Sending unit 70 can also receive an RTSP teardown request to terminate the streaming session; in response to this RTSP teardown request, sending unit 70 can stop transmitting media data for the corresponding session to client device 40.
[0040] Similarly, receiving unit 52 can initiate a media stream by initially sending an RTSP description request to server device 60. The RTSP description request can indicate the data types supported by client device 40. Then, receiving unit 52 can receive from server device 60 a reply specifying an available media stream (such as media content 64) that can be transmitted to client device 40 along with a corresponding network location identifier (such as a Uniform Resource Locator (URL) or Uniform Resource Name (URN)).
[0041] Then, receiving unit 52 can generate an RTSP establishment request and transmit it to server device 60. As noted above, the RTSP establishment request may include a network location identifier for the requested media data (e.g., media content 64) and a transport specification, such as the local port used to receive RTP data and control data (e.g., RTCP data) on client device 40. In response, receiving unit 52 may receive an acknowledgment from server device 60, including the port of server device 60 used to transmit media data and control data.
[0042] After a media session is established between server device 60 and client device 40, the transmitting unit 70 of server device 60 can transmit media data (e.g., media data packets) to client device 40 according to the media session. Server device 60 and client device 40 can exchange control data (e.g., RTCP data) indicating, for example, the reception statistics of client device 40, so that server device 60 can perform congestion control or otherwise diagnose and resolve transmission failures.
[0043] Network interface 54 can receive selected media and provide it to receiving unit 52, which in turn can provide the media data to decapsulation unit 50. Decapsulation unit 50 can decapsulate the elements of a video file into a PES stream, unpack the PES stream to retrieve encoded data, and transmit the encoded data to audio decoder 46 or video decoder 48, depending on whether the encoded data is part of an audio stream or a video stream (e.g., as indicated by the PES packet header of the stream). Audio decoder 46 decodes the encoded audio data and transmits the decoded audio data to audio output 42, while video decoder 48 decodes the encoded video data and transmits the decoded video data (which may include multiple views of the stream) to video output 44.
[0044] According to the technology disclosed herein, video decoder 48 can be configured to apply one or more video decoding neural networks to video data. Specifically, video decoder 48 may receive data defining the neural networks, for example, in or outside the video bitstream. For each neural network, video decoder 48 may receive data representing the purpose or type of the neural network, one or more structures of the neural network, weight values of the neural network, bias values of the neural network, and / or identifiers (e.g., sequence numbers) of the neural network.
[0045] Initially, the video decoder 48 can receive an encoded video dataset, as well as data indicating which neural network in the neural network should be applied to the encoded video data or to which part of the video data. For example, the data specifying the neural networks can indicate the sequential arrangement of the neural networks that process the video data in order.
[0046] According to the technology disclosed herein, the video decoder 48 can also receive updates to one or more neural networks in the neural network. Such updates can indicate whether to use an alternative neural network for a specific task, whether to skip processing of a specific task by the neural network, whether to discard one or more internal droptable structures from the neural network, or other changes to the neural network, such as changes to bias values, weight values, configuration parameters, etc. Therefore, the video decoder 48 can apply updates to the corresponding neural network and then decode the video data using the appropriately updated neural network. In some examples, the video decoder 48 can retain a copy of a previous version of the neural network associated with a corresponding sequence number, allowing the reuse of the previous version of the neural network, as identified by the corresponding sequence number.
[0047] In some examples, video decoder 48 may receive data representing a neural network and / or updates to the neural network from the video bitstream itself. In some examples, video decoder 48 may receive either or both of the data representing a neural network and / or updates to the neural network from, for example, receiving unit 52, network interface 54, or other external components. For example, receiving unit 52 may be configured to determine whether an RTP packet includes data representing a neural network and / or updates to the neural network in, for example, the RTP packet payload or the RTP packet extended header. Where the RTP packet payload includes data representing a neural network and / or updates to the neural network, receiving unit 52 may determine whether a given RTP packet includes such neural network data by first examining the header of the RTP packet, which may indicate whether the payload includes such neural network data.
[0048] Thus, according to the technology disclosed herein, the video decoder 48 can be dynamically reconfigured to apply different neural network sets and / or reconfigure the neural networks.
[0049] The video encoder 28, video decoder 48, audio encoder 26, audio decoder 46, encapsulation unit 30, receiving unit 52, and decapsulation unit 50 can all be implemented as any of a variety of suitable processing circuits, such as one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic circuits, software, hardware, firmware, or any combination thereof. Each of the video encoder 28 and video decoder 48 can be included in one or more encoders or decoders, and either the video encoder or the video decoder can be integrated as part of a combined video encoder / decoder (CODEC). Similarly, each of the audio encoder 26 and audio decoder 46 can be included in one or more encoders or decoders, and either the audio encoder or the audio decoder can be integrated as part of a combined CODEC. The apparatus including the video encoder 28, video decoder 48, audio encoder 26, audio decoder 46, encapsulation unit 30, receiving unit 52, and / or decapsulation unit 50 can include integrated circuits, microprocessors, and / or wireless communication devices, such as cellular phones.
[0050] Client device 40, server device 60, and / or content preparation device 20 may be configured to operate according to the techniques of this disclosure. For illustrative purposes, these techniques are described with respect to client device 40 and server device 60. However, it should be understood that content preparation device 20 may also be configured to perform these techniques as an alternative to (or other than) server device 60.
[0051] Encapsulation unit 30 can form NAL units, which include a header identifying the program to which the NAL unit belongs and a payload, such as audio data, video data, or data describing the transport or program stream corresponding to the NAL unit. For example, in H.264 / AVC, a NAL unit includes a 1-byte header and a payload of varying size. NAL units whose payloads include video data can include video data at various granularities. For example, a NAL unit can include video data blocks, multiple blocks, video data slices, or entire frames of video data. Encapsulation unit 30 can receive encoded video data in PES packet format with elementary streams from video encoder 28. Encapsulation unit 30 can associate each elementary stream with its corresponding program.
[0052] The encapsulation unit 30 can also assemble access units based on multiple NAL units. Generally, an access unit may include one or more NAL units representing a video data frame, and the corresponding audio data (when the audio data is available). Access units typically include all NAL units for a single output time instance, such as all audio and video data for a single time instance. For example, if each view has a frame rate of 20 frames per second (fps), each time instance may correspond to a time interval of 0.05 seconds. During this time interval, specific frames for all views with the same access unit (same time instance) can be rendered simultaneously. In one example, an access unit may include a decoded image within a time instance, which can be rendered as the primary decoded image.
[0053] Therefore, an access unit can include all audio and video frames of a common time instance, such as all views corresponding to time X. This disclosure also refers to the encoded picture of a particular view as a "view component." That is, a view component can include pictures (or frames) encoded for a particular view at a particular time. Therefore, an access unit can be defined as including all view components of a common time instance. The decoding order of the access units need not be the same as the output or display order.
[0054] After the encapsulation unit 30 has assembled the NAL units and / or access units into a video file based on the received data, the encapsulation unit 30 passes the video file to the output interface 32 for output. In some examples, the encapsulation unit 30 may store the video file locally or transmit the video file to a remote server via the output interface 32, instead of transmitting the video file directly to the client device 40. The output interface 32 may include, for example, a transmitter, a transceiver, a device for writing data to a computer-readable medium (such as, for example, an optical drive, a magnetic media drive (e.g., a floppy disk drive)), a universal serial bus (USB) port, a network interface, or other output interface. The output interface 32 outputs the video file to a computer-readable medium, such as, for example, a transmitting signal, a magnetic medium, an optical medium, a memory, a flash drive, or other computer-readable media.
[0055] Network interface 54 can receive NAL units or access units via network 74, and provide NAL units or access units to decapsulation unit 50 via receiving unit 52. Decapsulation unit 50 can decapsulate the elements of the video file into a PES stream, unpack the PES stream to retrieve encoded data, and transmit the encoded data to audio decoder 46 or video decoder 48, depending on whether the encoded data is part of an audio stream or a video stream (e.g., as indicated by the PES packet header of the stream). Audio decoder 46 decodes the encoded audio data and transmits the decoded audio data to audio output 42, while video decoder 48 decodes the encoded video data and transmits the decoded video data (which may include multiple views of the stream) to video output 44.
[0056] Figure 2 This is a conceptual diagram illustrating an example system capable of implementing the techniques of this disclosure. In this example, Figure 2 A system comprising encoder 102, bitrate result 106, decoder 104, task network 112, and task result 114 is described. Generally, encoder 102 receives input video / image data 100, which may represent a single image, image sequence, video data, etc. Encoder 102 encodes the received data according to a defined set of neural network-based coding tasks. Encoder 102 can test various neural networks and / or neural network configurations to determine (e.g., based on rate-distortion metrics) which neural network / configuration produces optimal performance. PSNR / SSIM result 110 and bitrate result 106 may include metrics that encoder 102 can use to determine the optimal neural network and / or neural network configuration data for a specific media dataset to be encoded.
[0057] According to the technology disclosed herein, encoder 102 can signal data representing neural network and / or configuration data in the media bitstream, in the packet header of packets in the media bitstream, for reception by decoder 104. Initially, encoder 102 can directly signal the neural network, wherein later in a media communication session, encoder 102 can signal updates to the previously signaled neural network.
[0058] Encoder 102 can test various modifications to the neural network, such as changes to quantization parameters. The neural network can include approximately one million weights and bias values, and encoder 102 can modify such values and encode data representing these modifications. Generally, the update values may be small. Therefore, differential signaling can reduce the bit rate associated with the neural network signaling updates.
[0059] Generally, encoder 102 can signal decoder 104 to the first data structure indicating an update to the second data structure. The second data structure can carry the neural network configuration associated with the compressed video (or image) bitstream. The compressed video bitstream can be used for machine learning tasks, such as machine-oriented video decoding (VCM), including object detection, object tracking, instance segmentation, etc.
[0060] The second data structure may include any or all of the following information: the type or purpose of the neural network, data defining the neural network structure, the weights and biases of the neural network, and / or identifiers, such as sequence numbers. The type or purpose of the neural network may be, for example, whether the neural network performs region-of-interest (ROI) based decoding, intra-frame decoding based on the neural network, frame-level spatial resampling, temporal resampling, or post-filtering. The neural network structure can be expressed in a standardized format, such as Open Neural Network Exchange (ONNX) or Neural Network Exchange Format (NNEF). The type or purpose of the neural network can uniquely identify the neural network. In some examples, the data representing the neural network structure can directly signal the structure (e.g., ONNX or NNEF data). In some examples, the data representing the neural network structure can signal the network location from which the neural network structure is retrieved, for example, a Uniform Resource Locator (URL), Uniform Resource Identifier (URI), or Uniform Resource Name (URN). Sequence numbers can be used to identify each data structure / neural network of the same type in ascending order.
[0061] As indicated above, the first data structure that can indicate an update to the neural network (second data structure) may include data representing the type or purpose of the second data structure / neural network, the sequence number of the second data structure, structural changes to the second neural network (which may be in a compressed format), changes to the weights and / or biases of the neural network (which may be in a compressed format), and / or new values or changes to other configuration parameters (e.g., quantization parameters, spatial downsampling ratio, or temporal downsampling ratio).
[0062] Encoder 102 can transmit data representing an initial neural network as High-Level Syntax (HLS) data in an image or video bitstream. HLS data can include, for example, Video Parameter Set (VPS), Sequence Parameter Set (SPS), Picture Parameter Set (PPS), Adaptation Parameter Set (APS), or Supplemental Augmentation Information (SEI) messages. Encoder 102 can determine which message type to use based on the frequency of changes to the neural network. For example, VPS can be used if the neural network is not expected to change during a session. SPS can be used if the neural network is expected to change between scenes or otherwise remain applicable to a frame sequence. PSP can be used to signal frequent changes, such as changes applicable to a specific picture.
[0063] Alternatively, as noted above, data representing neural networks and / or changes to neural networks can be signaled out of band, for example, via a transmission control protocol (TCP) (e.g., as a file transfer), or via a real-time transport protocol (RTP) or a unified datagram protocol (UDP), or via HTTP over QUIC and UDP, or via RTP over QUIC and UDP. In some examples, two data structures can be carried in the RTP packet payload, and the payload type specified in the RTP packet header can indicate VCM metadata. In some examples, two data structures can be carried in an RTP packet header extension included in the RTP packet, which also carries data of a bitstream encoded by encoder 102. In other examples, data structures can be signaled from encoder 102 to decoder 104 via a data channel, wherein the data channel can operate on a flow control transmission protocol (SCTP) for Web Real-Time Communication (WebRTC) or IP Multimedia Subsystem (IMS) communication.
[0064] Decoder 104 can then receive data for the neural network and updates to the neural network. Decoder 104 can apply the updates to the neural network and then use the updated neural network to decode the received media data. This produces reconstructed video / image 108. Task network 112 can perform one or more of a variety of neural network-based processing tasks on the reconstructed video / image 108, such as for object recognition, autonomous driving / robotics control, or other machine-based tasks. Task result 114 represents the result of such task processing.
[0065] Figure 3 This is a diagram illustrating an example computational diagram according to the technology of this disclosure. In some examples, the technology of this disclosure includes identifiers of structures that can be modified, as discussed above. Figure 3 The examples depict the structure, including convolution (conv) unit 182, ReLU unit 184, MaxPool unit 190, convolution unit 194, ReLU unit 198, convolution unit 204, convolution unit 206, global average pooling unit 212, and softmax unit 216.
[0066] In this example, convolutional unit 182 receives data_0 180. Convolutional unit 182 can have a W of <64 × 3 × 3 × 3> and <64> B. Convolutional unit 182 can output conv1_1 184, and ReLU unit 186 can process conv1_1 184 to form conv1_2 188. MaxPool unit 190 can process conv1_2 188 to form pool1_1 192. Convolutional unit 194 can have W <16 × 64 × 1 × 1> and <16> B. Convolutional unit 194 can process pool1_1 192 to form fire2 / squeeze1x1_1 196.
[0067] ReLU unit 198 can process fire2 / squeeze1x1_1 196 to form fire2 / squeeze1x1_2 200 and / or fire2 / squeeze1x1_2 202 or both. Convolution unit 204 can have W <64 × 16 × 1 × 1> and <64> The B, and can process fire2 / squeeze1x1_2 200 to form fire2 / expand1x1_1 208. The convolutional unit 206 can have a W of <64 × 16 × 3 × 3> and <64> The B unit can process fire2 / squeeze1x1_2 202 to form fire2 / expand3x3_1 210. Additional units, not shown in Figure 5, can form part of the task network. Finally, the global averaging pool unit 212 can receive the processed data and form pool10_1 214. The Softmax unit 216 can process pool10_1 214 to form softamxout_1 218.
[0068] The task network repository can include data representing various identifiers that indicate the structure of the computation graph, such as Figure 3 As shown. The computation graph can be, for example, an MLP, CNN, LSTM, GRU, transformer, etc. The task network store can transfer this data to the entities involved in executing the task network, such as user equipment (UE). For example, the UE can perform tasks related to... Figure 2 The functionality described in Task Network 112.
[0069] The task network repository can also transmit computation graph description tools or instructions for computation graph description tools. Examples of computation graph description tools include the Open Neural Network Exchange (ONNX) and the Neural Network Exchange Format (NNEF). In some examples, the task network repository may optionally transmit URIs, URLs, or URNs defining and marking structural modifications, parameter configuration changes, etc.
[0070] Figure 4 This is a flowchart illustrating an example method for updating a neural network according to the technology of this disclosure. Figure 4 The method can be performed by a network device (such as a cloud-based server) that receives video data and transmits it to a neural network for processing, for example, to perform machine-oriented video decoding (VCM). Figure 1 Video decoder 48 or Figure 2 The decoder 104 can be configured to perform Figure 4 The method. For explanatory purposes, for Figure 2 Decoder 104 is used to interpret Figure 4 The method.
[0071] Initially, decoder 104 receives data (250) representing multiple neural networks of various types. That is, each neural network can have an associated type. The type of neural network can represent the processing task performed by that neural network. Such processing tasks can include, for example, region-of-interest (ROI) based decoding, intra-frame prediction decoding, frame-level spatial resampling, temporal resampling, post-filtering, or other such processing tasks. Multiple neural networks can be associated with a specific video bitstream.
[0072] Decoder 104 can receive encoded video data from a video bitstream and provide the encoded video data to multiple neural networks to decode the video bitstream. Additionally or alternatively, the decoded video data can be sent to some or all of the neural networks to perform various VCM tasks, such as object recognition, object detection, object tracking, and instance segmentation.
[0073] At any given time, according to the techniques of this disclosure, it may be necessary to update one or more neural networks within a neural network. The update may be applied only for a specified time period, or the update may be applied continuously (unless the neural network is updated again later). Therefore, decoder 104 may receive an update for one of the neural networks, and the update may specify the type of neural network to which the update is applied (252). In some examples, the update may specify an identifier for the neural network version, such as a sequence number.
[0074] Updates can be included in the high-level syntax (HLS) of the video data. For example, updates can be included in a Supplemental Enhancement Information (SEI) message or a parameter set such as a Sequence Parameter Set (SPS) or Picture Parameter Set (PPS). Additionally or alternatively, neural network updates can be specified in the packet header (such as an RTP or RTSP header extension). For example, the header can indicate that the RTP / RTSP packet payload includes neural network update data, or the neural network update data can be included in the header itself. As another example, the update data can indicate the location from which data to be used to update the neural network is retrieved, such as a URL, URI, or URN of the data to be used to update the neural network. As yet another example, decoder 104 can initiate a connection (e.g., IMS or WebRTC) via a data channel specified in the update data to retrieve the neural network update data.
[0075] Although Figure 4The text only indicates a single update, but update data can include updates for one or more neural networks out of a set of multiple neural networks. An update can represent changes to, for example, the structure of the neural network, the weights within the network, the biases to be applied, parameter values, etc. Generally, data representing an update can indicate differences (or increments) relative to an existing neural network, rather than specifying a completely new replacement neural network.
[0076] Decoder 104 can determine a neural network (254) that matches the type specified in the update among multiple neural networks. Decoder 104 can then update the determined neural network (256). Thus, upon receiving new video data (258), decoder 104 can provide the video data to the neural network (260), including the updated neural network, and receive processed data (262) based on the video data.
[0077] so, Figure 4 The method representation includes examples of the following methods: receiving data representing multiple neural networks associated with a video bitstream, each of the multiple neural networks having a different type; receiving data representing an update to at least one of the neural networks, the data including a type corresponding to the at least one neural network and a neural network structure for the update; updating the neural networks according to the data representing the update to generate an updated neural network; and providing video data from the video bitstream to the updated neural network so that the updated neural network processes the video data.
[0078] Various examples of the techniques disclosed herein are summarized in the following provisions:
[0079] Clause 1: A method for processing video data, the method comprising: receiving data representing a neural network associated with a video bitstream; receiving data representing an update to the neural network; updating the neural network based on the data representing the update to generate an updated neural network; and providing video data from the video bitstream to the updated neural network to cause the updated neural network to process the video data.
[0080] Clause 2: The method according to Clause 1, wherein the video bitstream includes an encoded video bitstream, the method further comprising decoding the video bitstream to form decoded video data, wherein providing the video data to the updated neural network includes providing the decoded video data to the updated neural network.
[0081] Clause 3: The method according to any one of Clauses 1 and 2, wherein the data representing the neural network includes one or more of the following: the type or purpose of the neural network, the structure of the neural network, the weight values of the neural network, the bias values of the neural network, or the identifier of the neural network.
[0082] Clause 4: The method according to Clause 3, wherein the data representing the neural network includes the type or purpose of the neural network, the type or purpose including at least one of region of interest (ROI) based decoding, neural network-based intra-frame prediction decoding, frame-level spatial resampling, temporal resampling, or post-filtering.
[0083] Clause 5: The method according to any one of Clauses 3 and 4, wherein the data representing the neural network includes: the structure of the neural network expressed in one of the Open Neural Network Exchange (ONNX) format, Neural Network Exchange Format (NNEF) or a format defined in a Uniform Resource Locator (URL), Uniform Resource Identifier (URI) or Uniform Resource Name (URN).
[0084] Clause 6: The method according to any one of Clauses 3 to 5, wherein the data representing the neural network includes the identifier, and wherein the identifier includes a sequence number.
[0085] Clause 7: The method according to any one of Clauses 1 to 6, wherein the data representing the update includes one or more of the following: the type of the neural network, the identifier of the neural network, the structural change of the neural network, the weight change of the neural network, the bias change of the neural network, a new value of the configuration parameter, or a change of the configuration parameter.
[0086] Clause 8: The configuration parameters described in Clause 7 include one of a quantization parameter, a spatial downsampling ratio, or a temporal downsampling ratio.
[0087] Clause 9: The method according to any one of Clauses 1 to 8, wherein receiving the data representing the update comprises: receiving the data representing the update in at least one of a Video Parameter Set (VPS), a Sequence Parameter Set (SPS), a Picture Parameter Set (PPS), or in a Supplemental Enhancement Information (SEI) message.
[0088] Clause 10: The method according to any one of Clauses 1 to 8, wherein receiving the data representing the update comprises: receiving the data representing the update via a Transmission Control Protocol (TCP), a Real-Time Transport Protocol (RTP), or a Uniform Datagram Protocol (UDP).
[0089] Clause 11: The method according to Clause 10, wherein receiving the data representing the update comprises: receiving the data representing the update in the payload of an RTP packet, the method further comprising extracting data from the header of the RTP packet, the header indicating that the payload includes the data representing the update.
[0090] Clause 12: The method according to Clause 10, wherein receiving the data representing the update comprises: receiving the data representing the update in an RTP packet header extension.
[0091] Clause 13: The method according to any one of Clauses 1 to 8, wherein receiving the data representing the update comprises: receiving the data representing the update in a data channel of either an IP Multimedia Subsystem (IMS) or Web Real-Time Communications (WebRTC).
[0092] Clause 14: An apparatus for processing video data, the apparatus comprising one or more components for performing a method according to any one of Clauses 1 to 13.
[0093] Clause 15: The device according to Clause 14, wherein the one or more components include a processing system, the processing system including one or more processors implemented in a circuit.
[0094] Clause 16: The device according to any one of Clauses 14 and 15 further includes a memory configured to store the video data.
[0095] Clause 17: The device according to any one of Clauses 14 to 16, the device further includes a display configured to display the decoded video data.
[0096] Clause 18: The device pursuant to any one of Clauses 14 to 17, wherein the device includes one or more of the following: a camera, a computer, a mobile device, a broadcast receiver device, or a set-top box.
[0097] Clause 19: An apparatus for processing video data, the apparatus comprising: means for receiving data representing a neural network associated with a video bitstream; means for receiving data representing an update to the neural network; means for updating the neural network according to the data representing the update to generate an updated neural network; and means for providing video data from the video bitstream to the updated neural network to enable the updated neural network to process the video data.
[0098] Clause 20: A method for processing video data, the method comprising: receiving data representing a plurality of neural networks associated with a video bitstream, each of the plurality of neural networks having a different type; receiving data representing an update to at least one of the neural networks, the data including a type corresponding to the at least one neural network and a neural network structure for the update; updating the neural networks according to the data representing the update to generate an updated neural network; and providing video data from the video bitstream to the updated neural network to cause the updated neural network to process the video data.
[0099] Clause 21: The method according to Clause 20, wherein the type represents a task corresponding to the updated data, and wherein each neural network in the neural network is configured to perform the corresponding task.
[0100] Clause 22: The method according to Clause 21 further includes determining at least one neural network in the neural network that performs the task indicated in the data representing the update.
[0101] Clause 23: The method according to Clause 21, wherein the plurality of neural networks are configured to perform at least one of the following: region of interest (ROI) based decoding, neural network-based intra-frame prediction decoding, frame-level spatial resampling, temporal resampling, or post-filtering.
[0102] Clause 24: The method according to Clause 20, wherein the video bitstream includes an encoded video bitstream, the method further comprising decoding the video bitstream to form decoded video data, wherein providing the video data to the updated neural network includes providing the decoded video data to the updated neural network.
[0103] Clause 25: The method according to Clause 20, wherein the video bitstream comprises an encoded video bitstream, and wherein providing the video data to the updated neural network comprises providing the encoded video data to the updated neural network.
[0104] Clause 26: The method according to Clause 20, wherein the data representing the neural network includes one or more weight values of the neural network, bias values of the neural network, or an identifier of the neural network.
[0105] Clause 27: The method according to Clause 26, wherein the data representing the neural network includes the identifier, and wherein the identifier includes a sequence number.
[0106] Clause 28: The method according to Clause 20, wherein the data representing the plurality of neural networks includes: the structure of the plurality of neural networks expressed in one of the Open Neural Network Exchange (ONNX) format, Neural Network Exchange Format (NNEF) or a format defined in a Uniform Resource Locator (URL), Uniform Resource Identifier (URI) or Uniform Resource Name (URN).
[0107] Clause 29: The method according to Clause 20, wherein the data indicating the update includes one or more of the following: an identifier of the at least one neural network in the neural network, a structural change of the at least one neural network in the neural network, a weight change of the at least one neural network in the neural network, a bias change of the at least one neural network in the neural network, a new value of a configuration parameter, or a change of the configuration parameter.
[0108] Clause 30: The method described in Clause 29, wherein the configuration parameter includes one of a quantization parameter, a spatial downsampling ratio, or a temporal downsampling ratio.
[0109] Clause 31: The method according to Clause 20, wherein receiving the data representing the update comprises: receiving the data representing the update in at least one of a Video Parameter Set (VPS), a Sequence Parameter Set (SPS), a Picture Parameter Set (PPS), or in a Supplemental Enhancement Information (SEI) message.
[0110] Clause 32: The method according to Clause 20, wherein receiving the data representing the update comprises: receiving the data representing the update via Transmission Control Protocol (TCP), Real-Time Transport Protocol (RTP), or User Datagram Protocol (UDP).
[0111] Clause 33: The method described in Clause 32 further includes: performing a session initiation for TCP, RTP, or UDP.
[0112] Clause 34: The method according to Clause 32, wherein receiving the data representing the update comprises: receiving the data representing the update in the payload of an RTP packet, the method further comprising extracting data from the header of the RTP packet, the header indicating that the payload includes the data representing the update.
[0113] Clause 35: The method according to Clause 32, wherein receiving the data representing the update comprises: receiving the data representing the update in an RTP packet header extension.
[0114] Clause 36: The method according to any one of Clauses 1 to 30, wherein receiving the data representing the update comprises: receiving the data representing the update in a data channel of either an IP Multimedia Subsystem (IMS) or Web Real-Time Communications (WebRTC).
[0115] Clause 37: An apparatus for processing video data, the apparatus comprising: a memory configured to store video data; and a processing system implemented in circuitry, the processing system being configured to: receive data representing a plurality of neural networks associated with a video bitstream, each of the plurality of neural networks having a different type; receive data representing an update to at least one of the neural networks, the data including a type corresponding to the at least one neural network and a neural network structure for the update; update the neural networks according to the data representing the update to generate an updated neural network; and provide video data from the video bitstream to the updated neural network to cause the updated neural network to process the video data.
[0116] Clause 38: The device according to Clause 37, wherein the type represents a task corresponding to the updated data, and wherein each neural network in the neural network is configured to perform the corresponding task, wherein the processing system is further configured to determine at least one neural network in the neural network that performs the task indicated in the updated data.
[0117] Clause 39: An apparatus for processing video data, the apparatus comprising: means for receiving data representing a plurality of neural networks associated with a video bitstream, each of the plurality of neural networks having a different type; means for receiving data representing an update to at least one of the neural networks, the data including a type corresponding to the at least one neural network and a neural network structure for the update; means for updating the neural networks according to the data representing the update to generate an updated neural network; and means for providing video data from the video bitstream to the updated neural network to cause the updated neural network to process the video data.
[0118] Clause 40: A method for processing video data, the method comprising: receiving data representing a plurality of neural networks associated with a video bitstream, each of the plurality of neural networks having a different type; receiving data representing an update to at least one of the neural networks, the data including a type corresponding to the at least one neural network and a neural network structure for the update; updating the neural networks according to the data representing the update to generate an updated neural network; and providing video data from the video bitstream to the updated neural network to cause the updated neural network to process the video data.
[0119] Clause 41: The method according to Clause 40, wherein the type represents a task corresponding to the data representing the update, and wherein each neural network in the neural network is configured to perform the corresponding task.
[0120] Clause 42: The method according to Clause 41 further includes determining at least one neural network in the neural network that performs the task indicated in the data representing the update.
[0121] Clause 43: The method according to any one of Clauses 41 and 42, wherein the plurality of neural networks are configured to perform at least one of the following: region of interest (ROI) based decoding, neural network-based intra-frame prediction decoding, frame-level spatial resampling, temporal resampling, or post-filtering.
[0122] Clause 44: The method according to any one of Clauses 40 to 43, wherein the video bitstream comprises an encoded video bitstream, the method further comprising decoding the video bitstream to form decoded video data, wherein providing the video data to the updated neural network comprises providing the decoded video data to the updated neural network.
[0123] Clause 45: The method according to any one of Clauses 40 to 43, wherein the video bitstream comprises an encoded video bitstream, and wherein providing the video data to the updated neural network comprises providing the encoded video data to the updated neural network.
[0124] Clause 46: The method according to any one of Clauses 40 to 45, wherein the data representing the neural network includes one or more weight values of the neural network, bias values of the neural network, or an identifier of the neural network.
[0125] Clause 47: The method according to Clause 46, wherein the data representing the neural network includes the identifier, and wherein the identifier includes a sequence number.
[0126] Clause 48: The method according to any one of Clauses 40 to 47, wherein the data representing the plurality of neural networks includes: the structure of the plurality of neural networks expressed in one of the Open Neural Network Exchange (ONNX) format, Neural Network Exchange Format (NNEF) or a format defined in a Uniform Resource Locator (URL), Uniform Resource Identifier (URI) or Uniform Resource Name (URN).
[0127] Clause 49: The method according to any one of Clauses 40 to 48, wherein the data representing the update includes one or more of the following: an identifier of the at least one neural network in the neural network, a structural change of the at least one neural network in the neural network, a weight change of the at least one neural network in the neural network, a bias change of the at least one neural network in the neural network, a new value of a configuration parameter, or a change of the configuration parameter.
[0128] Clause 50: The method according to Clause 49, wherein the configuration parameter includes one of a quantization parameter, a spatial downsampling ratio, or a temporal downsampling ratio.
[0129] Clause 51: The method according to any one of Clauses 40 to 50, wherein receiving the data representing the update comprises: receiving the data representing the update in at least one of a Video Parameter Set (VPS), a Sequence Parameter Set (SPS), a Picture Parameter Set (PPS), or in a Supplemental Enhancement Information (SEI) message.
[0130] Clause 52: The method according to any one of Clauses 40 to 52, wherein receiving the data representing the update comprises: receiving the data representing the update via Transmission Control Protocol (TCP), Real-Time Transport Protocol (RTP), or User Datagram Protocol (UDP).
[0131] Clause 53: The method described in Clause 52 further includes: performing a session initiation for TCP, RTP, or UDP.
[0132] Clause 54: The method according to any one of Clauses 52 and 53, wherein receiving the data representing the update comprises: receiving the data representing the update in the payload of an RTP packet, the method further comprising extracting data from the header of the RTP packet, the header indicating that the payload includes the data representing the update.
[0133] Clause 55: The method according to any one of Clauses 52 to 54, wherein receiving the data representing the update comprises: receiving the data representing the update in an RTP packet header extension.
[0134] Clause 56: The method according to any one of Clauses 40 to 55, wherein receiving the data representing the update comprises: receiving the data representing the update in a data channel of either an IP Multimedia Subsystem (IMS) or Web Real-Time Communications (WebRTC).
[0135] Clause 57: An apparatus comprising one or more components for performing the method according to any one of Clauses 40 to 56.
[0136] Clause 58: A computer-readable storage medium having instructions stored thereon, which, when executed, cause a processor to perform a method according to any one of Clauses 40 to 56.
[0137] In one or more examples, the described functionality may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functionality may be stored as one or more instructions or code on a computer-readable medium or transmitted via a computer-readable medium and executed by a hardware-based processing unit. A computer-readable medium may include a computer-readable storage medium (which corresponds to a tangible medium such as a data storage medium) or a communication medium, including, for example, any medium that facilitates the transfer of a computer program from one place to another according to a communication protocol. Thus, a computer-readable medium may generally correspond to (1) a non-transitory tangible computer-readable storage medium, or (2) a communication medium such as a signal or carrier wave. A data storage medium may be any available medium that can be accessed by one or more computers or one or more processors to extract instructions, code, and / or data structures for implementing the techniques described in this disclosure. Computer program products may include computer-readable media.
[0138] By way of example, and not limitation, such computer-readable storage media may include RAM, ROM, EEPROM, CD-ROM or other optical disc storage devices, magnetic disk storage devices or other magnetic storage devices, flash memory, or any other medium capable of storing desired program code in the form of instructions or data structures and accessible by a computer. Furthermore, any connection is appropriately referred to as a computer-readable medium. For example, if instructions are transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies (such as infrared, radio, and microwave), then coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies (such as infrared, radio, and microwave) are included in the definition of medium. However, it should be understood that computer-readable storage media and data storage media do not include connections, carrier waves, signals, or other transient media, but instead refer to non-transient tangible storage media. As used herein, disks and optical discs include: compact optical discs (CDs), laser optical discs, optical discs, digital versatile optical discs (DVDs), floppy disks, and Blu-ray discs, wherein disks typically reproduce data magnetically, while optical discs reproduce data optically using lasers. The combinations described above should also be included within the scope of computer-readable media.
[0139] Instructions can be executed by one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable arrays (FPGAs), or other equivalent integrated or discrete logic circuits. Therefore, as used herein, the term "processor" can refer to any of the foregoing structures or any other structure suitable for implementing the techniques described herein. Additionally, in some aspects, the functionality described herein can be provided within dedicated hardware and / or software modules configured for encoding and decoding, or incorporated into combined codecs. Furthermore, these techniques can be fully implemented in one or more circuit or logic elements.
[0140] The techniques disclosed herein can be implemented in a wide variety of devices or apparatuses, including wireless mobile phones, integrated circuits (ICs), or IC sets (e.g., chipsets). Various components, modules, or units are described in this disclosure to emphasize functional aspects of a device configured to perform the disclosed techniques, but implementation by different hardware units is not necessarily required. Specifically, as described above, various units may be combined within a codec hardware unit, or may be provided by a collection of interoperable hardware units (including one or more processors as described above) combined with appropriate software and / or firmware.
[0141] Various examples have been described. These and other examples are within the scope of the following claims.
Claims
1. A method for processing video data, the method comprising: Receive data representing multiple neural networks associated with a video bitstream, each of the multiple neural networks having a different type; Receive data representing an update to at least one of the neural networks, the data including a type corresponding to the at least one neural network and a neural network structure for the update; The neural network is updated based on the data representing the update to generate an updated neural network; as well as Video data from the video bitstream is fed into the updated neural network so that the updated neural network processes the video data.
2. The method of claim 1, wherein the type represents a task corresponding to the updated data, and wherein each neural network in the neural network is configured to perform the corresponding task.
3. The method of claim 2, further comprising determining at least one neural network in the neural network that performs the task indicated in the data representing the update.
4. The method of claim 2, wherein the plurality of neural networks are configured to perform at least one of the following: region of interest (ROI) based decoding, neural network-based intra-frame prediction decoding, frame-level spatial resampling, temporal resampling, or post-filtering.
5. The method of claim 1, wherein the video bitstream comprises an encoded video bitstream, the method further comprising decoding the video bitstream to form decoded video data, wherein providing the video data to the updated neural network comprises providing the decoded video data to the updated neural network.
6. The method of claim 1, wherein the video bitstream comprises an encoded video bitstream, and wherein providing the video data to the updated neural network comprises providing the encoded video bitstream to the updated neural network.
7. The method of claim 1, wherein the data representing the neural network includes one or more weight values of the neural network, bias values of the neural network, or an identifier of the neural network.
8. The method of claim 7, wherein the data representing the neural network includes the identifier, and wherein the identifier includes a sequence number.
9. The method of claim 1, wherein the data representing the plurality of neural networks comprises: The structures of the plurality of neural networks are expressed in one of the following formats: Open Neural Network Exchange (ONNX) format, Neural Network Exchange (NNEF) format, or a format defined in a Uniform Resource Locator (URL), Uniform Resource Identifier (URI), or Uniform Resource Name (URN).
10. The method of claim 1, wherein the data representing the update includes one or more of the following: an identifier of the at least one neural network in the neural network, a structural change of the at least one neural network in the neural network, a weight change of the at least one neural network in the neural network, a bias change of the at least one neural network in the neural network, a new value of a configuration parameter, or a change of the configuration parameter.
11. The method of claim 10, wherein the configuration parameters include one of a quantization parameter, a spatial downsampling ratio, or a temporal downsampling ratio.
12. The method of claim 1, wherein receiving the data representing the update comprises: The data representing the update is received in at least one of the Video Parameter Set (VPS), Sequence Parameter Set (SPS), and Picture Parameter Set (PPS) or in a Supplemental Enhancement Information (SEI) message.
13. The method of claim 1, wherein receiving the data representing the update comprises: The data representing the update is received via Transmit Control Protocol (TCP), Real-Time Transport Protocol (RTP), or User Datagram Protocol (UDP).
14. The method according to claim 13, further comprising: Performs session initiation for TCP, RTP, or UDP.
15. The method of claim 13, wherein receiving the data representing the update comprises: The method further includes receiving data representing an update in the payload of an RTP packet, the method also including extracting data from the header of the RTP packet, the header indicating that the payload includes the data representing the update.
16. The method of claim 13, wherein receiving the data representing the update comprises: The data indicating an update is received in the RTP packet header extension.
17. The method of claim 1, wherein receiving the data representing the update comprises: The data representing the update is received in the data channel of either the IP Multimedia Subsystem (IMS) or Web Real-Time Communication (WebRTC).
18. An apparatus for processing video data, the apparatus comprising: A memory configured to store video data; and A processing system implemented in a circuit, the processing system being configured to: Receive data representing multiple neural networks associated with a video bitstream, each of the multiple neural networks having a different type; Receive data representing an update to at least one of the neural networks, the data including a type corresponding to the at least one neural network and a neural network structure for the update; The neural network is updated based on the data representing the update to generate an updated neural network; as well as Video data from the video bitstream is fed into the updated neural network so that the updated neural network processes the video data.
19. The apparatus of claim 18, wherein the type represents a task corresponding to the updated data, and wherein each neural network in the neural network is configured to perform the corresponding task, wherein the processing system is further configured to determine at least one neural network in the neural network that performs the task indicated in the updated data.
20. An apparatus for processing video data, the apparatus comprising: A component for receiving data representing multiple neural networks associated with a video bitstream, each of the multiple neural networks being of a different type; A component for receiving data representing an update to at least one of the neural networks, the data including a type corresponding to the at least one neural network and a neural network structure for the update; Components for updating the neural network based on the data representing the update to generate an updated neural network; and Components for providing video data from the video bitstream to the updated neural network so that the updated neural network processes the video data.