Compressing videos for feedforward view synthesis

ROI encoding using neural networks to dynamically adjust compression based on regions of interest addresses bandwidth challenges in video streaming, ensuring seamless playback across different connection speeds.

WO2025160413A1PCT designated stage Publication Date: 2025-07-31GOOGLE LLC
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/012972
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-24
Filing Date
2025-01-24
Publication Date
2025-07-31

AI Technical Summary

Technical Problem

Existing video streaming pipelines require large bandwidth and can result in time delays and interrupted displays due to the need to stream multiple video streams from different camera perspectives, especially in low-bitrate connections.

Method used

Implementing region-of-interest (ROI) encoding by compressing video frames based on the likelihood of including a region of interest, using a neural network to generate ROI maps and adjust quantization parameters dynamically, thereby reducing the number of bits required for non-ROI areas.

Benefits of technology

This approach reduces bandwidth requirements, enabling smooth streaming on various connection speeds and improving user experience by minimizing time delays and interruptions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025012972_31072025_PF_FP_ABST
    Figure US2025012972_31072025_PF_FP_ABST
Patent Text Reader

Abstract

A method including receiving a plurality of video streams, generating a plurality of frames of the plurality of video streams, generating a plurality of region of interest maps associated with the plurality of frames, and compressing the plurality of frames as a plurality of compressed frames including quantizing the plurality of frames using a quantization parameter associated with a corresponding one of the plurality of region of interest maps.
Need to check novelty before this filing date? Find Prior Art

Description

COMPRESSING VIDEOS FOR FEEDFORWARDVIEW SYNTHESISCROSS-REFERENCE TO RELATED APPLICATION

[0001] This application claims priority to U.S. Provisional Patent Application No. 63 / 624,729, filed on January 24, 2024, the disclosure of which is incorporated by reference herein in its entirety.FIELD

[0002] Implementations relate to synthesizing an image or frame of a video from two or more images or frames from two or more videos.BACKGROUND

[0003] Image and / or video synthesis can include generating one image and / or one video based on multiple images or videos. Novel view synthesis can use multiple images or videos taken from different view perspectives as input and use neural algorithms to interpolate the view perspectives associated with multiple images or videos into a novel (or new) view perspective.SUMMARY

[0004] Implementations relate to streaming multiple real-time videos for synthesizing a video from the multiple streamed videos. The videos are compressed to reduce the size of the video streams. In addition, the size of the video streams can be further reduced in size by reducing the number of bits used to represent portions of the video (e.g., video frames) that do not include a region of interest.

[0005] In a general aspect, a device, a system, a non-transitory computer- readable medium (having stored thereon computer executable program code which can be executed on a computer system), and / or a method can perform a process with a method including receiving a plurality of video streams, generating a plurality of frames of the plurality of video streams, generating a plurality of region of interest maps associated with the plurality of frames, and compressing the plurality of frames as a plurality of compressed frames including quantizing the plurality of frames using aquantization parameter associated with a corresponding one of the plurality of region of interest maps.BRIEF DESCRIPTION OF THE DRAWINGS

[0006] Example implementations will become more fully understood from the detailed description given herein below and the accompanying drawings, wherein like elements are represented by like reference numerals, which are given by way of illustration only and thus are not limiting of the example implementations.

[0007] FIG. 1 illustrates a novel view or view perspective image generation according to an example implementation.

[0008] FIG. 2 illustrates a data flow (or pipeline) associated with streaming and generating a video according to an example implementation.

[0009] FIG. 3 illustrates a data flow associated with encoding a video according to at least one example implementation.

[0010] FIG. 4 illustrates a block diagram of a data flow of a model architecture showing a data flow for a training and inference procedure according to at least one example implementation.

[0011] FIG. 5 illustrates a block diagram of a design for a differentiable codec proxy according to an example implementation.

[0012] FIG. 6 is a block diagram of a method of compressing two or more video streams according to an example implementation.

[0013] FIG. 7 is a block diagram of a method of training a model according to an example implementation.

[0014] FIG. 8 is a block diagram of a method of streaming a video according to an example implementation.

[0015] It should be noted that these Figures are intended to illustrate the general characteristics of methods, and / or structures utilized in certain example implementations and to supplement the written description provided below. These drawings are not, however, to scale and may not precisely reflect the precise structural or performance characteristics of any given implementation and should not be interpreted as defining or limiting the range of values or properties encompassed by example implementations. For example, the positioning of modules and / or structural elements may be reduced or exaggerated for clarity. The use of similar or identicalreference numbers in the various drawings is intended to indicate the presence of a similar or identical element or feature.DETAILED DESCRIPTION

[0016] Novel view synthesis can include generating or synthesizing a single video using video captured by multiple cameras. This single video can be displayed (or rendered for display) on a display. The video is generated having a novel (e.g., new) perspective in respect to some content of the video. The novel view perspective is not the same (or substantially equivalent to) any of the view perspectives associated with the multiple cameras used to capture the video. Accordingly, generating the single video can include generating the single video based on the video captured by the multiple cameras and with the novel view perspective.

[0017] Novel view synthesis can include multiple learned solutions that provide high quality, photorealistic results. Novel view synthesis can use multiple images or videos taken from different view perspectives as input and use neural or conventional structure-from-motion algorithms to interpolate the view perspectives associated with multiple images or videos into a novel (or new) view perspective. The quality afforded by these approaches is suitable for their use in real-world applications including, for example, re-rendering an environment with different climate, text-to-3D asset creation, simultaneous localization and mapping (SLAM), and the like. Some application opportunities for novel view synthesis include real-time streaming. For example, live streaming events for real-time free-viewpoint video, 3D telepresence to replace 2D video conferencing, photorealistic 3D cloud gaming, robotics applications, and the like.

[0018] At least one technical problem with existing video streaming pipelines is that streaming multiple video streams associated with multiple cameras can require a large amount of bandwidth. Further, real-time view synthesis that uses multiple video streams to generate a real-time video can be too slow to provide a desirable experience. For example, requiring a large amount of bandwidth to stream real-time video often results in time delays and / or interrupted display of video. In addition, some network connections may not have sufficient bandwidth (e.g., a low-bitrate connection) for streaming multiple real-time videos. At least one technical solution can include using region-of-interest (ROI) encoding. ROI encoding can include compressing portions of a frame of a video to include a different number of bits for different portions of the video. For example, a portion(s) of a frame of a video having a high likelihood ofincluding a ROI can be compressed with more bits, a first number of bits, and / or a higher number of bits and a portion(s) of the frame of a video having a low likelihood of including a ROI can be compressed with fewer bits, a second number of bits and / or lower number of bits. In other words, the number of bits associated with a frame of video can be reduced by including fewer bits (e.g., a lower bitrate) for portions of the frame of a video that don’t include a ROI. At least one technical effect of the technical solution can be to improve a user experience by reducing the number of bits associated with a streaming video. Reducing the number of bits (e.g., a lower bitrate) can, for example, enable streaming on any bandwidth and / or low-bitrate connection. Reducing the number of bits (e.g., a lower bitrate) can improve a user’s experience by, for example, preventing time delays, queuing, stoppage, and / or interrupted display of video.

[0019] FIG. 1 illustrates a novel (e.g., new or not captured using a camera) view or view perspective image generation or image synthesis example according to an example implementation. As shown in FIG. 1, cameras (e.g., two or more cameras, a plurality of cameras, and the like) Cl, C2, C3, C4, C5, can be used (e.g., each can be used) to capture an image and / or a video (e.g., frames of a video). The video can be of a scene, an object, an event, a person(s), and / or the like). For example, FIG. 1 illustrates cameras Cl, C2, C3, C4, C5 being used to capture a video of a person 105. For example, cameras Cl, C2, C3, C4, C5 can be used to capture a video of person 105 in a video conference.

[0020] Cameras Cl, C2, C3, C4, C5 can capture video from a view perspective (sometimes referred to as a view). For example, camera Cl can have view perspective Pl, camera C2 can have view perspective P2, camera C3 can have view perspective P3, camera C4 can have view perspective P4, and camera C5 can have view perspective P5. In the example of FIG. 1, five (5) cameras having respective view perspectives are illustrated. However, any number of camera(s) having a respective view perspective(s) can be used and are within the scope of this disclosure.

[0021] In some implementations, during playback of the video captured by cameras Cl, C2, C3, C4, C5, a single video can be generated or synthesized using the video captured by cameras Cl, C2, C3, C4, C5. This single video can be displayed (or rendered for display) on display 115. The video is generated having a view perspective P6 in respect to user 110 (the viewer of the video). However, as shown in FIG. 1, view perspective P6 is not the same (or substantially equivalent to) any of the viewperspectives Pl, P2, P3, P4, P5. Therefore, view perspective P6 is a new or novel view perspective. Accordingly, generating the video (including person 105’) can include generating the video with view perspective P6.

[0022] In some implementations, the video captured by cameras Cl, C2, C3, C4, C5 can be streamed from a local device to a remote device (e.g., respective computer devices). In some implementations, the video captured by cameras Cl, C2, C3, C4, C5 can be streamed as individual video streams. Arrow 120 represents streaming (e.g., as individual streams) the video captured by cameras Cl, C2, C3, C4, C5 from a local device to a remote device.

[0023] FIG. 2 illustrates a data flow (or pipeline) associated with streaming and generating a video according to an example implementation. As shown in FIG. 2, the data flow includes an encoder 205, a transmitter 210, a receiver 215, a decoder 220, an image generator 225 and the display 115. In some implementations, encoder 205 can be configured to compress the video captured by cameras Cl, C2, C3, C4, C5 (e.g., compress each of the video captured by cameras Cl, C2, C3, C4, C5 individually). Encoder 205 can use any compression scheme or codec. For example, encoder 205 can use a high efficiency video coding (HVEC) video codec or standard. Decoder 220 can be configured to perform the inverse of encoder 205. In other words, decoder 220 can be configured to decompress the compressed video. In some implementations, the decoder 220 can be configured to decompress individual streams. In some implementations, the decoder 220 can be configured to generate reconstructed video representing the video captured by cameras Cl, C2, C3, C4, C5.

[0024] Transmitter 210 and receiver 215 together provide the functionality to stream video. In some implementations, the transmitter 210 can be an element of a local device and the receiver 215 can be an element of a remote device. In some implementations, the transmitter 210 can be an element of a first edge node in a network and the receiver 215 can be an element of a second edge node in the network. In some implementations, transmitter 210 can be configured to generate packet(s) including video to be communicated using a (e.g., wired or wireless) communications standard. In some implementations, receiver 215 can be configured to unpack video from the packet(s) including video.

[0025] Image generator 225 can be configured to generate an image, a video, a frame of the video, and / or the like based on a plurality of images, videos, frames of the videos, and / or the like. Image generation is sometimes referred to as image synthesis.Video generation is sometimes referred to as video synthesis. Video frame generation is sometimes referred to as video frame synthesis. In some implementations, image generator 225 can be configured to generate video based on reconstructed video representing the video captured by cameras Cl, C2, C3, C4, C5. In some implementations, image generator 225 can be configured to generate video having a new or novel view perspective. In some implementations, image generator 225 can be configured to generate video having a new or novel view perspective based on view perspective 230.

[0026] Video compression for real-time view synthesis remains a relatively underexplored area compared to some video streaming applications, such as cloud gaming, multi-view codecs, and red-green-blue-depth (RGBD) streaming. Given that many view-synthesis techniques integrate information from multiple views of the same subject, there is potential to reduce the overall bitrate of the stream through region-of- interest (ROI) encoding. Some implementations described herein include a differentiable video proxy for, for example, the HEVC video codec to enable gradients from a feedforward view synthesis method to backpropagate into a preprocessor neural network which computes ROI maps. To that effect, some implementations jointly train the view synthesis network and the preprocessor network to optimize a combined image reconstruction and bitrate loss objective. Some methods achieve significant bitrate savings by learning to compress regions or views that provide lesser contributions to the reconstruction network. A further reduction in bitrate can be achieved by training the joint system to perform background replacement for rendered views. For example, operating at 40-60 frames per second graphical processing unit (GPU), this preprocessor network can attain a significant reduction in bitrate compared to previous baselines, with minimal impact on quality.

[0027] While some high-quality view synthesis methods can be too slow for real-time inference or capture, there are some approaches that are almost fast enough for commercial use. Unfortunately, display end-points, such as smartphones or standalone virtual reality (VR) headsets, lack the compute resources to be able to run some of these algorithms. Therefore, some implementations can use a streaming cloud architecture. For example, referring to FIG. 2, the receiver 215, the decoder 220, and the image generator 225 can be performed using a streaming cloud architecture. In this case, another encoder, transmitter, receiver, decoder pipeline (not shown) would be included between the image generator 225 and the display 115.

[0028] In such an architecture, some implementations can be configured to stream multiple (e.g. 4-8) high resolution streams from an input camera rig to a display device and / or a host device in the cloud. At higher resolutions (e.g. 4K), this could amount to over 200Mbps of upstream bandwidth using standard video codecs which would be impractically expensive for a large proportion of locations. Fortunately, these input streams represent different views of the same object. Therefore, this bandwidth can be reduced by exploiting redundancy between views. Moreover, depending on the specific view synthesis method, regions of the input views that are not visible from the target may not be sent over the network, further reducing bandwidth.

[0029] Some implementations present a technique to reduce upstream bandwidth for a generalizable view synthesis method. To do this, some implementations described herein jointly train a neural ROI preprocessor network and a view-synthesis network. Some implementations are trained end-to-end from a lightweight preprocessor network, through the compression algorithm, and to the outputs of novel view synthesis, and backpropagate gradients back to the preprocessor. In some implementations, this allows the preprocessor network to learn which regions of the upstream images provide useful information to the view synthesis network and instruct the codec to heavily compress unimportant regions. In some implementations, these regions tend to be redundant between streams or less necessary for high-quality texture synthesis from the novel view synthesis network. In some implementations, the potential bandwidth savings are even greater if only the foreground needs to be displayed. Some implementations achieve end-to-end training by incorporating a differentiable codec proxy that mimics the performance of the real codec. During inference, some implementations replace the differentiable codec proxy with the real codec hardware.

[0030] In some implementations, this joint compress-and-reconstruct training procedure is a shift away from hand-crafted compression procedures intended to minimize distortion in the inputs of view synthesis. Some system implementations can be configured to learn to compress video for view synthesis, and can be optimized to minimize the distortion incurred in the rendered output of the view synthesis method while maintaining a high, user-configurable compression ratio.

[0031] FIG. 3 illustrates a data flow associated with encoding a video according to at least one example implementation. As shown in FIG. 3, encoder 205 can include a preprocessor 305 block, a blocking 310 block, a transform 315 block, and aquantization 320 block. The blocking 310, the transform 315, and the quantization 320 can be associated with a codec 325.

[0032] The preprocessor 305 can be configured to generate an ROI map. The preprocessor 305 can be a model configured to generate an ROI map. The preprocessor 305 can be a neural network configured to generate an ROI map. In some implementations, the preprocessor network can be trained to learn which regions of an image provide useful information or data to an image generator (e.g., image generator 225) or a view synthesis network. The preprocessor 305 can be configured to cause the codec 325 (e.g., an HEVC codec) to heavily compress unimportant regions. In some implementations, these regions can be redundant between streams or less necessary for high-quality image synthesis by an image generator (e.g., image generator 225) or a view synthesis network.

[0033] In some implementations, preprocessor 305 can be configured to generate an ROI map that codec 325 uses to determine quantization parameter (QP) values for quantization 320. In some implementations, preprocessor 305 can be configured to operate on all of the provided reference RGB frames from the input cameras as well as reference (e.g., the perspective view of a camera whose images are used in the novel view synthesis or Pl, P2, P3, P4, and P5 in FIG. l) and target (e.g., the novel view to be synthesized) camera poses. In some implementations, the inclusion of camera poses can reduce the number of bits that are allocated to regions that are not used by the view synthesis algorithm, simply because those regions are not on any epipolar line from the target camera. In some implementations, at least one of the objectives of the preprocessor 305 can be to detect matching regions between different camera views (e.g., similar to stereo correspondence). In some implementations, these correspondences can help the preprocessor 305 to assign fewer bits to foreground regions that are repeated between (or included in) multiple views, acting as a kind of visual de-duplication for view synthesis. Thus, preprocessor 305 should facilitate this kind of stereo correspondence operation, either by directly aligning reference view features across views using an epipolar constraint, or indirectly through attention. Attention mechanisms can be used to refine intermediate feature maps, which can improve the realism of the generated images.

[0034] As shown in FIG. 3, preprocessor 305 can be configured to receive the video captured by cameras Cl, C2, C3, C4, C5, pass thru the video captured by cameras Cl, C2, C3, C4, C5, and generate an ROI map Ml, M2, M3, M4, M5 associated with arespective video captured by cameras Cl, C2, C3, C4, C5. In some implementations, the preprocessor 305 generates an ROI map for a frame(s) of a respective video captured by cameras Cl, C2, C3, C4, C5. A ROI map can be visualized as a heat map, an attention map, a data map, and / or the like. A heat map can depict values for a main variable of interest across two axis variables as a grid of pixels (e.g., color or greyscale pixels). An attention map can show the relative importance of each region relative to the other regions in an image. An ROI map can have the same structure as an image, video or video frame. For example, if a video frame is 1000 x 1000 pixels, the ROI map is also 1000 x 1000 pixels (or some other variable). However, with the ROI map each pixel (or some other variable) represents an interest level. For example, visually, a darker pixel can represent a pixel having more interest and a lighter pixel can represent a pixel having less interest.

[0035] A ROI map can be a data structure including a quantization parameter (QP) associated with a portion (as small as one pixel) of a video or frame of video. In some implementations, a video or video frame can include a plurality of portions and each of the plurality of portions can have an associated QP in the ROI map. In some implementations, the ROI map can include QP values that are spatially distributed throughout the frame. For example, the ROI map can have a first QP associated with a first portion of the video or frame and a second QP associated with a second portion of the video or frame. In some implementations, the ROI map can have a first QP associated with a first two or more portions of the video or frame and a second QP associated with a second two or more portions of the video or frame.

[0036] In some implementations, a quantization parameter (QP) determines the step size for associating the transformed coefficients with a finite set of steps. Large values of QP represent big steps that crudely approximate the spatial transform, so that most of the signal can be captured by only a few coefficients. A large QP value can result in more compression, larger quantization steps, and lower quality. A small QP value results in less compression, smaller quantization steps, and more detail. In the context of the ROI map visualized as a heat map, an attention map, a data map, and / or the like, a dark pixel can have a relatively small QP value and a light pixel can have a relatively large QP value. Alternatively, a dark pixel can have a relatively large QP value and a light pixel can have a relatively small QP value.

[0037] Blocking 310 (sometimes called partitioning) can be configured to partition an image, a video, a frame of a video and the like into a plurality of portions(sometimes referred to as a unit, a block or block of pixels). In some implementations, blocking 310 can partition using a division into square coding tree blocks (CTB). The blocks can have different sizes (e.g., number of pixels). In some implementations, the CTB can be defined across both luma and chroma and each CTB can be per-channel (with chroma subsampling the Cb / Cr channels are half size). For example, the sizes can range from 4x4, 8x8, 16x16, 16x32, 32x32, 32x64, 64x64 and the like luma samples.

[0038] Transform 315 can be configured to transform each block into a set of coefficients. Transform 315 can be a discrete cosine transform (DCT), a discrete sine transform (DST), and the like. Quantization 320 can be configured to compress the coefficients. Quantization 320 can be a lossy operation when compressing an image, video, or video frame. Quantization 320 can compress the coefficients using (or based on) the QP value (using ROI map Ml, M2, M3, M4, M5). Briefly, in a video compression operation, each video frame is divided into smaller blocks called macroblocks. The macroblocks are transformed generating a set of coefficients. A quantization matrix is applied to the coefficients of the macroblock based on a spatial frequency content. The resulting values are rounded off based on a QP. In some implementations, QP is a dimensionless value that can range from 0 to 51, where 0 represents lossless encoding and 51 represents a maximum or coarse compression. In some implementations, QP is a dimensionless value that can range from 0 to 255, where 0 represents lossless encoding and 255 represents a maximum or coarse compression. Other minimums and maximum QP values are within the scope of this disclosure. Further, some examples may prevent a QP value at the minimum or maximum from being used. For example, the minimum of 0 may not be used. Instead an alternate minimum value (e.g., 4, 6, 8, and the like) may be used.

[0039] Detail view 330 illustrates a portion of the video captured by camera C 1.Detail view 330 illustrates a number of blocks associated with a portion of a frame of video captured by camera Cl. The blocks or macroblocks can be transformed and quantized. As shown in FIG. 3, a QP is input to quantization 320 and the QP is associated with an ROI map. If a block or macroblock is a region of interest, QP may be a low value (e.g., closer to 0, but often not 0). If a block or macroblock is not a region of interest, QP may be a high value (e.g., closer to 51).

[0040] Novel view synthesis is a classical problem where, given a set of images from different view perspectives of a scene, a novel view is computed that accuratelyreproduces motion parallax and visibility occlusions. At least some implementations can be trained on a large set of multi-view images of scenes and then run in a feedforward manner on a set of multi-view images for a new scene. This allows for increasingly faster scene reconstruction times and sparser input views.

[0041] As mentioned above, preprocessor 305 can be a model and / or a neural network. In addition, for example, some implementations can be configured to work with any generalizable neural approach to novel view synthesis. Therefore, preprocessor 305 and / or image generator 225 can be trained to perform any of the techniques described herein.

[0042] By training end to end, the preprocessors (e.g., preprocessor 305) described herein learn which parts of the input images are useful to the view synthesis network and which parts aren’t and so can tune the inputs to the video codec accordingly. In order to present some implementations within a working system, some implementations use Generalizable Patch-Based Neural Rendering (GPNR) as a view synthesis network, but some implementations should work with any differentiable approach to novel view synthesis.

[0043] FIG. 4 illustrates a block diagram of a data flow of a model architecture according to at least one example implementation. As shown in FIG. 4, a data flow of a model architecture can include preprocessor 305, codec 325, decoder 220, image generator 225, a multiplier sampler 415 block, and ground truth block. Codec 325 can include an HEVC 405 block and a differentiable proxy 410 block.

[0044] Assuming an A-camera capturing system, some implementations can compress the N video streams in realtime for remote 3D reconstruction and view synthesis. Therefore, at a certain time t, the input to some implementations of a compressor can be defined as N pairs of datawhere I denotes synchronous image frames and denotes camera view direction, respectively. In some implementations, since a standard video codec partitions a frame into spatial blocks, one can dynamically adjust the compression level of each block by varying the corresponding QP value based on its contribution to the target view synthesis. In some implementations, the QP values of all the blocks assemble into what some implementations define as an ROI map M.

[0045] FIG. 4 depicts the overall data flow of some implementations of the algorithm. There are at least three components: preprocessor 305, codec 325 (e.g., ahardware-accelerated video codec) and image generator 225 (sometimes referred to as a view synthesizer). Taking the multiview dataanovel target viewand multiplier sampler 415 (e.g., a desired rate-distortion multiplier) as input, preprocessor 305 E can be configured to predict an ROI map Mn for each frame In. Codec 325 then utilizes Mn to encode In and decode into?;1, which are leveraged by image generator 225 (e.g., a view synthesizer) G to render into the target view 425 image C.

[0046] Below, the architecture of the preprocessor 305 and image generator 225 (e.g., a view synthesizer) are described according to some implementations. During inference time, these components are used in conjunction with, for example, codec 325 as a standard video codec. In addition, how the HEVC codec can be replaced with a differentiable proxy in order to train all of these components end-to-end according to some implementations is described.

[0047] Preprocessor 305 can be configured to produce the ROI maps Ml, M2, M3, ..., Mn that codec 325 uses to determine QP values. Preprocessor 305 can operate on all of the provided reference RGB frames from the input cameras as well as, for example, reference and target camera poses. In some implementations, the inclusion of camera poses can reduce the number of bits that are allocated to regions that are not used by the novel view synthesis algorithm, because those regions are not on any epipolar line from the target camera. In some implementations, at least one of the objectives of the preprocessor 305 can be to detect matching regions between different camera views, in a similar fashion to stereo correspondence.

[0048] In some implementations, these correspondences help the preprocessor 305 assign fewer bits to foreground regions that are repeated between views, acting as a visual de-duplication for view synthesis. Thus, the architecture of this component should facilitate this kind of stereo correspondence operation, either by directly aligning reference view features across views using an epipolar constraint, or indirectly through attention.

[0049] Some implementations chose to use a plane sweep volume (PSV) to explicitly align reference view features for this purpose. In some implementations, the preprocessor initially uses a 2D feature extraction network on the input views at multiple scales. In some implementations, these features are scaled to the appropriate resolution then resampled onto a PSV which is aligned on the target view. In someimplementations, this operation implicitly culls features that are not visible from the target view.

[0050] In some implementations, after applying convolutional blocks on this PSV, these features can be warped into the input views and accumulated. In some implementations, to do so, a number of points on each ray are sampled on all of the input views, resampled from the target PSV, then accumulated via a single convolution operation across the ray.

[0051] In some implementations, after the ROI maps Ml, M2, M3, ..., Mn have been generated by the preprocessor 305 and the videos have been encoded (e.g., using codec 325) and decoded (using decoder 220), the resulting decoded views and camera poses are passed into the image generator 225. In some implementations, the image generator 225 can be a feedforward view synthesis model. In some implementations, a feedforward view synthesis model can be a Generalized Patch-Based Neural Rendering (GPNR) model, although any feed-forward view synthesis method could be used (e.g. PixelNeRF, IBRNet, etc.). For a target ray, GPNR first extract patch along the epipolar line. These extracted patches undergo transformation through a visual feature transformer, which accepts a sequence of patches corresponding to the same depth of projection as input. In some implementations, the cross-view feature aggregation that occurs in NeRF acts as a kind of denoiser.

[0052] In some implementations, the refined patch features are then processed by the epipolar transformer that predicts a weight for each epipolar point. In some implementations, these weights are used to aggregate the feature along the epipolar line which are then fed into the view transformer to predict a weight for each reference view. In some implementations, the weights for the epipolar points and the reference views are used to blend the interpolated RGB value at the epipolar points to predict the color for the target ray. In some implementations, during training, a number of random rays (e.g., 1024) can be sampled from the target view frustum, rendered, and compared against a target view. Some implementations use an LI loss as a reconstruction loss and also include the weight regularization term in GPNR.

[0053] Some implementations can be configured to synthesize comparatively clean reconstructions from input frames that exhibit heavy compression artifacts. This behavior can be partially explained by GPNR using some patches as shape cues rather than sources for texture synthesis, but it can also be explained by the principle that aggregating multiple patches tends to smooth and denoise the inputs. In someimplementations, a change to GPNR can include omitting the world coordinate input to the view transformer.

[0054] FIG. 4 also illustrates a block diagram of a data flow of a model architecture showing a data flow for a training and inference procedure according to at least one example implementation. In some implementations, to enable gradient back- propagation and end-to-end learning, a differentiable proxy 410 or special differentiable proxy of the codec can be in place during training. Some implementations adopt a block-based differentiable codec proxy (e.g., codec 325 or a modification of codec 325). The block-based differentiable codec proxy can simulate the DCT and quantization operations within an industrial video codec. The block-based differentiable codec proxy can be configured to generate two differentiable outputs, a bitrate and a decoded image. This codec proxy does not attempt to simulate motion compensation, which is a practical simplification that some implementations make so that the model does not need to carry forward state across frames. Refinement operations, such as the in-loop filtering operations used in HEVC, can also be omitted from the codec proxy, for simplicity.

[0055] In some implementations, the proxy first partitions the image into a set of 32x32 blocks ( / = {bi}). Each block is transformed via the DCT into a set of coefficients (b = {xj}) which are then quantized to be compressed. In HEVC, the quantization dividend (e.g., qstep A) is determined from the quantization parameter QP by the following relationshipThe quantization in an example proxy is then made differentiable using a straight- through quantizer for a DCT coefficient xj,where [ • ] is the rounding operation, W is the true rounding noise, andrepresents the stop gradient operator. Note dQ(xj) / dxj is nonzero across nearly the entirety of its domain. This can enable gradients to propagate back through the codec proxy andobtain a decoded output I (shown as I’ in FIG. 4) which is differentiable with respect to the input image I.

[0056] The above operations can occur in the Rec. 709 YCbCr colorspace. To simplify the codec design, chroma subsampling is not performed (e.g., YUV444). This process is visualized in FIG. 5. In some implementations, a rate proxy is employed to provide a differentiable output, defined aswhere the estimated bitrate of the image I is computed across multiple blocks {bi} and thus multiple DCT coefficients {xj} quantized with qstep Ai. The log function is a replacement of the original nondifferentiable indicator function, a is a scaling factor between the estimated and true rates, computed for each frame as below:Then the bitrate loss across multiple views can be described as

[0057] Some implementations use the view synthesizerfor estimating distortion via neural rendering. As mentioned, our choice of view synthesizer is GPNR. For brevity, some implementations useJra }> < >) to describe the output color of rendering a ray with GPNR. The distortion loss can then be described as:where 4> isground truth target view image (e.g., ground truth 420), and(- represents the pixel that is specified by the ray (p. A ground truth image can be an image used to compare the generated image to. For training efficiency, instead of rendering the full target image, rays are sampled at random. However, note that from the perspective of the gradient into the differentiable codec (i.e.only the pixels of which lie on the epipolar lines of the sampled ray (p sampled by GPNR have validgradients. This motivates a large batch size during training, to improve the weak signal- to-noise ratio from the sparse gradient.

[0058] In some implementations, the model described in previous sections can solve the problem of general-purpose view synthesis. This kind of approach might be useful for full-scene reconstruction, but the overall performance of the system can be improved if the view synthesis component is specialized for 3D foreground-background segmentation.

[0059] In some implementations, the overall system is optimized end-to-end to minimize both rate and distortion, but only the distortion of the foreground regions. In some implementations, the preprocessor network effectively creates a coarse segmentation of the input views and culls details in the inputs that are not necessary for the final reconstruction. In some implementations, the view synthesis module learns to only reconstruct the foreground on a ray-by-ray basis as well. In some implementations, notably, reference view masks are never used at any point during training or inference only target view foreground masks are used at all. In some implementations, the preprocessor effectively learns how to segment the reference views in order to provide the most useful information to the view synthesis module while also minimizing bitrate.

[0060] In some implementations, in addition to performing simple foregroundbackground segmentation, implementations can also be modified to perform coarse background reconstruction by using a background loss term that is weighted lower than the foreground. In some implementations, this is an intermediate between hard segmentation and not segmenting at all.

[0061] At least one alternative approach is to use the alpha channel of the rendered output of the model as a measure of foreground and then introduce a matting loss against the target matte. Some implementations modify GPNR to return an auxiliary alpha output (as it does not normally do alpha compositing) and introduce a final binary cross entropy auxiliary loss against this output.

[0062] In some implementations, a multiplier sampler can be configured to generate a Lagrange multiplier A can be used to control the trade-off between bitrate and distortion, such that the overall loss 430 can be defined as,£ ~ £« + * Zu (8) where:L is the overall loss,£<^is the bitrate loss,Z is the Lagrange multiply er, and£® is distortion loss.

[0063] Some learning-based methods train one model per multiplier value, which reduces flexibility and increases the model memory footprint for compression tasks. Some implementations instead randomly sample multiplier values from a truncated Gaussian distribution to specify this Gaussian during training. In some implementations, each multiplier value is passed into the preprocessor network. In some implementations, this allows the preprocessor network to learn the distribution of the Lagrange multiplier and enables the user to make a tradeoff between quality and bitrate at runtime.

[0064] FIG. 5 illustrates a block diagram of a data flow for a differentiable codec proxy according to an example implementation. As shown in FIG. 5. The data flow includes the blocking 310 block, the transform 315 block, an input frame 505 block, an ROI map 510 block, a rate proxy 515 block, an inverse transform 520 block, a deblocking 525 block, a decoded frame 530 block, and a quantization proxy 535 block.

[0065] The input frame 505 I is partitioned into blocks by blocking 310. Each block can have an associated qstep Az which is used by the quantization proxy and the rate proxy. R(xj , Ai ) is the estimated bitrate of transform 315 (e.g., a DCT) coefficient xj. I can be the result of emulating encode and decode using this process. In some implementations, the input video frame can be transformed into a 4:4:4 YCbCr image beforehand.

[0066] Quantization proxy 535 can represent both quantization and inverse quantization of a transformed image. Inverse transform 520 can perform the opposite of the transform 315, deblocking 525 can perform the opposite of blocking 310 and decoded frame 530 can represent a reconstructed version of input frame 505.

[0067] FIG. 6 is a block diagram of a method of compressing two or more video streams according to an example implementation. As shown in FIG. 6, in step S605 two or more video streams are received. For example, the two or more video streams can be associated with a live event. For example, the two or more streams can be associatedwith a video conference. In some implementations, the two or more streams can be captured by cameras having different view perspectives of the live event (e.g., video conference). In some implementations, the two or more streams can be received at substantially the same time. In some implementations, the two or more streams can be high resolution (e.g., 4K, and the like). In other words, the two or more streams can include a large number of pixels, bits, or bytes to represent each frame of the video. Therefore, each frame of the two or more video streams should be compressed before communicating the video from a first node in a network to a second node in a network. The second node can be a server or cloud-based node.

[0068] In step S610 two or more region of interest maps are generated based on a frame of the two or more video streams. For example, a region of interest can be a portion of a video frame that includes content that contributes to synthesizing an image based on compressing two or more video streams. For example, regions or views that contribute to the image synthesis can be regions of interest whereas regions or views that provide lesser contributions may not be regions of interest. For example, in a video conference the participant(s) may correspond to regions of interest and a background may not be a region of interest. Further, the face of a participant may be of more interest than the body of the participant. Some implementations may not be associated with a video conference. In these cases, focal objects could be of interest. For example, an athlete in a sporting event, a musician in a concert, an object (e.g., a car, a building, a tree, a flower, and / or the like), and / or the like). In some implementations a region of interest map can include two or more regions of interest. For example, a video conference can include two or more participants in the same location.

[0069] In some implementations, the region of interest (ROI) map can be generated or visualized as a heat map, an attention map, a data map, and / or the like. The ROI map can have the same structure as an image, video or video frame. For example, if a video frame is 1000 x 1000 pixels, the ROI map is also 1000 x 1000 pixels (or some other variable). However, with the ROI map each pixel (or some other variable) represents an interest level. For example, visually, a darker pixel can represent a pixel having more interest and a lighter pixel can represent a pixel having less interest. Further, numerically, a number representing black (e.g., RGB(0, 0, 0)) can represent more interest and a number representing white (e.g., RGB(255, 255, 255)) can represent a pixel having less interest (note that interest can have a range accordingly shades of grey can represent varying levels of interest).

[0070] In some implementations, the frames of the two or more video streams can be intended for use in synthesizing an image. Therefore, the image synthesizing technique (e.g., model and / or neural network) can be considered when generating a region of interest map. For example, a region of interest can be a portion of the frame where more information or data is needed to synthesize an image. Further, a portion of the frame where less information or data is needed to synthesize an image may not be a region of interest. For example, if there are five (5) video streams, the image synthesizing technique may only need minimal information from a portion of each of the five (5) frames to synthesize one frame for display. Therefore, the portion of each of the five (5) frames can be compressed more highly with minimal effect on the quality of the synthesized frame.

[0071] In step S615 the frame of the two or more video streams are transformed.For example, a frame of a video stream can be partitioned into portions, blocks, or macroblocks. Then, each portion, block, or macroblock can be transformed using, for example, a DCT, a DST, or the like.

[0072] In step S620 a quantization parameter is determined for a portion of the frame of the two or more video streams based on the region of interest map. In some implementations, quantization can be performed on each pixel, portion, block, or macroblock of the frame. For example, in the case where quantization is performed on each pixel, each pixel can have a corresponding quantization parameter. Further, if the region of interest (ROI) map can be visualized or generated as a heat map, an attention map, a data map, and / or the like, each pixel (or some other variable) of the ROI map can have a corresponding quantization parameter. For example, a look-up table (or some other data structure) can be used to look-up a quantization parameter based on a value of the region of interest map (as a heat map, an attention map, a data map, and / or the like) corresponding to the pixel being quantized. In some implementations, each variable of the region of interest map can be a quantization parameter. In some implementations, the quantization parameter can be based on a plurality of quantization parameters. For example, if a portion, block, or macroblock of the frame, are quantized with one quantization parameter, a plurality of quantization parameters can be averaged, a maximum can be determined, a minimum can be determined, a median can be determined, and the like.

[0073] In step S625 the portion of the frame of the two or more video streams is quantized using the quantization parameter. Quantization is a compression processwith loss. For example, Quantization can reduce the precision of a video frame (a portion of a video frame, a pixel of a video frame, and the like) by converting it to a lower-precision format. In some implementations, the quantization can be adaptive quantization that varies compression within a frame to distribute bits to provide more data to areas of a frame that are (or include) a region of interest and less data to areas of a frame that are not (or do not include) a region of interest.

[0074] In other words, adaptive quantization (e.g., varying the quantization parameter) can compress portions of a frame of a video to include a different number of bits for different portions of the video. For example, a portion(s) of a frame of a video having a high likelihood of including a ROI can be compressed to a first or higher number of bits and a portion(s) of the frame of a video having a low likelihood of including a ROI can be compressed to a second or lower number of bits. In other words, the number of bits associated with a frame of video can be reduced by including fewer bits (e.g., a lower bitrate) for portions of the frame of a video that don’t include a ROI.

[0075] FIG. 7 is a block diagram of a method of training a model according to an example implementation. As shown in FIG. 7, in step S705 two or more frames associated with video streams are received. For example, two or more frames can be training frames where there are known or previously identified regions of interest. Therefore, the video streams can be training video streams. In some implementations, the two or more frames can have (e.g., be captured by cameras having) different view perspectives. In some implementations, the two or more frames can be high resolution frames (e.g., 4K, high dynamic range (HDR) video, and the like). In other words, the two or more frames can include a large number of pixels, bits, or bytes to represent each frame.

[0076] In step S710 a target view perspective and a ground truth image are received. For example, the target view perspective can be a view perspective for an image generated based on the two or more frames and the ground truth image can be an image used to compare the generated image to. As an example, the two or more frames can include five (5) frames captured at substantially the same time, each frame having an associated view perspective (noting that the view perspectives should be different view perspectives). At substantially the same time, a sixth frame can be captured, the sixth frame can have a different view perspective as compared to the view perspective of the five (5) frames. In an example implementation, the sixth frame can be used as the ground truth image and the view perspective of the sixth frame can beused as the target view perspective.

[0077] In step S715 two or more region of interest maps are generated based on the two or more frames. For example, a region of interest can be a portion of a video frame that includes content that contributes to synthesizing an image based on compressing two or more video streams. For example, regions or views that contribute to the image synthesis can be regions of interest whereas regions or views that provide lesser contributions may not be regions of interest. For example, in a video conference the participant(s) may correspond to regions of interest and a background may not be a region of interest. Further, the face of a participant may be of more interest than the body of the participant. Some implementations may not be associated with a video conference. In these cases, focal objects could be of interest. For example, an athlete in a sporting event, a musician in a concert, an object (e.g., a car, a building, a tree, a flower, and / or the like)), and / or the like. In some implementations a region of interest map can include two or more regions of interest. For example, a video conference can include two or more participants in the same location.

[0078] In some implementations, the region of interest map can be generated or visualized as a heat map, an attention map, a data map, and / or the like. The ROI map can have the same structure as an image, video or video frame. For example, if a video frame is 1000 x 1000 pixels, the ROI map is also 1000 x 1000 pixels (or some other variable). However, with the ROI map each pixel (or some other variable) represents an interest level. For example, visually, a darker pixel can represent a pixel having more interest and a lighter pixel can represent a pixel having less interest. Further, numerically, a number representing black (e.g., RGB(0, 0, 0)) can represent more interest and a number representing white (e.g., RGB(255, 255, 255)) can represent a pixel having less interest (note that interest can have a range accordingly shades of grey can represent varying levels of interest).

[0079] In some implementations, the frames of the two or more video streams can be intended for us in synthesizing an image. Therefore, the image synthesizing technique (e.g., model and / or neural network) can be considered when generating a region of interest map. For example, a region of interest can be a portion of the frame where more information or data is needed to synthesize an image. Further, a portion of the frame where less information or data is needed to synthesize an image may not be a region of interest. For example, if there are five (5) video streams, the image synthesizing technique may only need minimal information from a portion of each ofthe five (5) frames to synthesize one frame for display. Therefore, the portion of each of the five (5) frames can be compressed more highly with minimal effect on the quality of the synthesized frame.

[0080] In step S720 the two or more frames are compressed based on the two or more region of interest maps using a differentiable proxy, and a quantization proxy. For example, the two or more frames can be partitioned, transformed and quantized. For example, the two or more frames can be partitioned into portions, blocks, or macroblocks. Then, each portion, block, or macroblock can be transformed using, for example, a DCT, a DST, or the like. As discussed above, the codec may use a differentiable proxy, and a quantization proxy (described in more detail above).

[0081] Next, a quantization parameter is determined for a portion of the frame of the two or more video streams based on the region of interest map. In some implementations, quantization can be performed on each pixel, portion, block, or macroblock of the frame. For example, in the case where quantization is performed on each pixel, each pixel can have a corresponding quantization parameter. Further, if the region of interest map can be visualized or generated as a heat map, an attention map, a data map, and / or the like, each pixel (or some other variable) of the ROI map can have a corresponding quantization parameter. For example, a look-up table (or some other data structure) can be used to look-up a quantization parameter based on a value of the region of interest map (as a heat map, an attention map, a data map, and / or the like) corresponding to the pixel being quantized. In some implementations, each variable of the region of interest map can be a quantization parameter. In some implementations, the quantization parameter can be based on a plurality of quantization parameters. For example, if a portion, block, or macroblock of the frame, are quantized with one quantization parameter, a plurality of quantization parameters can be averaged, a maximum can be determined, a minimum can be determined, a median can be determined, and the like.

[0082] Finally, the portion of the frame of the two or more frames is quantized using the quantization parameter. Quantization is a lossy compression process. For example, Quantization can reduce the precision of a video frame (a portion of a video frame, a pixel of a video frame, and the like) by converting it to a lower-precision format. In some implementations, the quantization can be adaptive quantization that varies compression within a frame to distribute bits to provide more data to areas of a frame that are (or include) a region of interest and less data to areas of a frame that arenot (or do not include) a region of interest.

[0083] In other words, adaptive quantization (e.g., varying the quantization parameter) can compress portions of a frame to include a different number of bits for different portions of the video. For example, a portion(s) of a frame having a high likelihood of including a ROI can be compressed to a first or higher number of bits and a portion(s) of the frame having a low likelihood of including a ROI can be compressed to a second or lower number of bits. In other words, the number of bits associated with a frame of video can be reduced by including fewer bits (e.g., a lower bitrate) for portions of the frame of a video that don’t include a ROI.

[0084] In step S725 the two or more frames are decompressed. For example, the two or more frames can be inverse quantized, inverse transformed, and deblocked. For brevity's sake, decompressing a frame can be the opposite (or inverse of) step S720 except there may be no determination of a quantization parameter based on a point of interest map. By contrast, a default or standard inverse quantization process may be used.

[0085] In step S730 an image is generated based on the decompressed two or more frames. In some implementations, the image is generated using a feedforward view synthesis model. In some implementations, a feedforward view synthesis model can be a Generalized Patch-Based Neural Rendering (GPNR) model, although any feed-forward view synthesis method could be used (e.g. PixelNeRF, IBRNet, etc.). For a target ray, GPNR first extracts patches along the epipolar line. These extracted patches undergo transformation through a visual feature transformer, which accepts a sequence of patches corresponding to the same depth of projection as input.

[0086] In some implementations, the refined patch features are then processed by the epipolar transformer that predicts a weight for each epipolar point. In some implementations, these weights are used to aggregate the feature along the epipolar line which are then fed into the view transformer to predict a weight for each reference view. In some implementations, the weights for the epipolar points and the reference views are used to blend the interpolated RGB value at the epipolar points to predict the color for the target ray. In some implementations, during training, a number of random rays (e.g., 1024) can be sampled from the target view frustum, rendered, and compared against a target view. Some implementations use an LI loss as a reconstruction loss, and also include the weight regularization term in GPNR.

[0087] Some implementations can be configured to synthesize comparatively clean reconstructions from input frames that exhibit heavy compression artifacts. This behavior can be partially explained by GPNR using some patches as shape cues rather than sources for texture synthesis, but it can also be explained by the principle that aggregating multiple patches tends to smooth and denoise the inputs.

[0088] In step S735 a loss is generated based on the image and the ground truth image. For example, the loss can be generated as described above with regard to Eqn. 8.

[0089] In step S740 a model is trained based on the loss. For example, the region of interest maps can be generated using a model (described in detail above regarding, for example, preprocessor 305). The model or neural network can include a plurality of weights. In an example implementation, the loss can be compared to criteria (e.g., a maximum and / or minimum loss). If the loss meets the criteria, the training process can end. If the loss does not meet the criteria, at least one of the weights associated with the model or neural network can be modified (e.g., increased and / or decreased) and the training process can be repeated beginning at step S705.

[0090] Example 1. FIG. 8 is a block diagram of a method of streaming a video according to an example implementation. As shown in FIG. 8, in step S805 receiving a plurality of video streams. In step S810 generating a plurality of frames of the plurality of video streams. In step S815 generating a plurality of region of interest maps associated with the plurality of frames. In step S820 compressing the plurality of frames as a plurality of compressed frames including quantizing the plurality of frames using a quantization parameter associated with a corresponding one of the plurality of region of interest maps.

[0091] Example 2. The method of Example 1 can further include streaming the plurality of compressed frames.

[0092] Example 3. The method of Example 1, wherein the plurality of video streams can have different view perspectives of a scene.

[0093] Example 4. The method of Example 1, wherein the generating of the plurality of region of interest maps can include using a model trained to indicate a redundancy between view perspectives associated with the plurality of frames.

[0094] Example 5. The method of Example 1, wherein the generating of the plurality of region of interest maps can include using a model trained to indicate a foreground and a background associated with the plurality of frames.

[0095] Example 6. The method of Example 1, wherein the generating of the plurality of region of interest maps can include using a model trained to indicate regions of the plurality of frames that include data necessary information for image synthesis.

[0096] Example 7. The method of Example 1, wherein the generating of the plurality of region of interest maps can include using a model trained to indicate regions of the plurality of frames that that are not used for image synthesis.

[0097] Example 8. The method of Example 1, wherein the plurality of region of interest maps can be heat maps, an attention map, a data map, and / or the like where a value of a pixel of the heat map, an attention map, a data map, and / or the like represents an interest level.

[0098] Example 9. The method of Example 8, wherein the quantization parameter can correspond to the interest level.

[0099] Example 10. The method of Example 8 can further include determining the quantization parameter based on a look-up table and the value of the pixel.

[0100] Example 11. The method of Example 1, wherein the generating of the plurality of region of interest maps can include using a model trained using a differential codec proxy, a quantization proxy, and a rate proxy.

[0101] Example 12. The method of Example 1, wherein the generating of the plurality of region of interest maps can include using a model trained using a loss based on a bitrate loss and a distortion loss.

[0102] Example 13. A method can include any combination of one or more of Example 1 to Example 12.

[0103] Example 14. A non-transitory computer-readable storage medium comprising instructions stored thereon that, when executed by at least one processor, are configured to cause a computing system to perform the method of any of Examples 1-13.

[0104] Example 15. An apparatus comprising means for performing the method of any of Examples 1-13.

[0105] Example 16. An apparatus comprising at least one processor and at least one memory including computer program code, the at least one memory and the computer program code configured to, with the at least one processor, cause the apparatus at least to perform the method of any of Examples 1-13.

[0106] Example implementations can include a non-transitory computer- readable storage medium comprising instructions stored thereon that, when executedby at least one processor, are configured to cause a computing system to perform any of the methods described above. Example implementations can include an apparatus including means for performing any of the methods described above. Example implementations can include an apparatus including at least one processor and at least one memory including computer program code, the at least one memory and the computer program code configured to, with the at least one processor, cause the apparatus at least to perform any of the methods described above.

[0107] Various implementations of the systems and techniques described here can be realized in digital electronic circuitry, integrated circuitry, specially designed ASICs (application specific integrated circuits), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which may be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.

[0108] These computer programs (also known as programs, software, software applications or code) include machine instructions for a programmable processor, and can be implemented in a high-level procedural and / or object-oriented programming language, and / or in assembly / machine language. As used herein, the terms “machine- readable medium” “computer-readable medium” refers to any computer program product, apparatus and / or device (e.g., magnetic discs, optical disks, memory, Programmable Logic Devices (PLDs)) used to provide machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term “machine-readable signal” refers to any signal used to provide machine instructions and / or data to a programmable processor.

[0109] To provide for interaction with a user, the systems and techniques described here can be implemented on a computer having a display device (a LED (light-emitting diode), or OLED (organic LED), or LCD (liquid crystal display) monitor / screen) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback(e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.

[0110] The systems and techniques described here can be implemented in a computing system that includes a back end component (e.g., as a data server), or that includes a middleware component (e.g., an application server), or that includes a front end component (e.g., a client computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described here), or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (“LAN”), a wide area network (“WAN”), and the Internet.

[0111] The computing system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other.

[0112] A number of implementations have been described. Nevertheless, it will be understood that various modifications may be made without departing from the spirit and scope of the specification.

[0113] In addition, the logic flows depicted in the figures do not require the particular order shown, or sequential order, to achieve desirable results. In addition, other steps may be provided, or steps may be eliminated, from the described flows, and other components may be added to, or removed from, the described systems. Accordingly, other implementations are within the scope of the following claims.

[0114] While certain features of the described implementations have been illustrated as described herein, many modifications, substitutions, changes and equivalents will now occur to those skilled in the art. It is, therefore, to be understood that the appended claims are intended to cover all such modifications and changes as fall within the scope of the implementations. It should be understood that they have been presented by way of example only, not limitation, and various changes in form and details may be made. Any portion of the apparatus and / or methods described herein may be combined in any combination, except mutually exclusive combinations. The implementations described herein can include various combinations and / or sub-combinations of the functions, components and / or features of the different implementations described.

[0115] While example implementations may include various modifications and alternative forms, implementations thereof are shown by way of example in the drawings and will herein be described in detail. It should be understood, however, that there is no intent to limit example implementations to the particular forms disclosed, but on the contrary, example implementations are to cover all modifications, equivalents, and alternatives falling within the scope of the claims. Like numbers refer to like elements throughout the description of the figures.

[0116] Some of the above example implementations are described as processes or methods depicted as flowcharts. Although the flowcharts describe the operations as sequential processes, many of the operations may be performed in parallel, concurrently or simultaneously. In addition, the order of operations may be re-arranged. The processes may be terminated when their operations are completed, but may also have additional steps not included in the figure. The processes may correspond to methods, functions, procedures, subroutines, subprograms, etc.

[0117] Methods discussed above, some of which are illustrated by the flow charts, may be implemented by hardware, software, firmware, middleware, microcode, hardware description languages, or any combination thereof. When implemented in software, firmware, middleware or microcode, the program code or code segments to perform the necessary tasks may be stored in a machine or computer readable medium such as a storage medium. A processor(s) may perform the necessary tasks.

[0118] Specific structural and functional details disclosed herein are merely representative for purposes of describing example implementations. Example implementations, however, be embodied in many alternate forms and should not be construed as limited to only the implementations set forth herein.

[0119] It will be understood that, although the terms first, second, etc. may be used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another. For example, a first element could be termed a second element, and, similarly, a second element could be termed a first element, without departing from the scope of example implementations. As used herein, the term and / or includes any and all combinations of one or more of the associated listed items.

[0120] It will be understood that when an element is referred to as being connected or coupled to another element, it can be directly connected or coupled to the other element or intervening elements may be present. In contrast, when an element is referred to as being directly connected or directly coupled to another element, there are no intervening elements present. Other words used to describe the relationship between elements should be interpreted in a like fashion (e.g., between versus directly between, adjacent versus directly adjacent, etc.).

[0121] The terminology used herein is for the purpose of describing particular implementations only and is not intended to be limiting of example implementations. As used herein, the singular forms a, an and the are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms comprises, comprising, includes and / or including, when used herein, specify the presence of stated features, integers, steps, operations, elements and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof.

[0122] It should also be noted that in some alternative implementations, the functions / acts noted may occur out of the order noted in the figures. For example, two figures shown in succession may in fact be executed concurrently or may sometimes be executed in the reverse order, depending upon the functionality / acts involved.

[0123] Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by one of ordinary skill in the art to which example implementations belong. It will be further understood that terms, e.g., those defined in commonly used dictionaries, should be interpreted as having a meaning that is consistent with their meaning in the context of the relevant art and will not be interpreted in an idealized or overly formal sense unless expressly so defined herein.

[0124] Portions of the above example implementations and corresponding detailed description are presented in terms of software, or algorithms and symbolic representations of operation on data bits within a computer memory. These descriptions and representations are the ones by which those of ordinary skill in the art effectively convey the substance of their work to others of ordinary skill in the art. An algorithm, as the term is used here, and as it is used generally, is conceived to be a self-consistent sequence of steps leading to a desired result. The steps are those requiring physicalmanipulations of physical quantities. Usually, though not necessarily, these quantities take the form of optical, electrical, or magnetic signals capable of being stored, transferred, combined, compared, and otherwise manipulated. It has proven convenient at times, principally for reasons of common usage, to refer to these signals as bits, values, elements, symbols, characters, terms, numbers, or the like.

[0125] In the above illustrative implementations, reference to acts and symbolic representations of operations (e.g., in the form of flowcharts) that may be implemented as program modules or functional processes include routines, programs, objects, components, data structures, etc., that perform particular tasks or implement particular abstract data types and may be described and / or implemented using existing hardware at existing structural elements. Such existing hardware may include one or more Central Processing Units (CPUs), digital signal processors (DSPs), applicationspecific-integrated-circuits, field programmable gate arrays (FPGAs) computers or the like.

[0126] It should be borne in mind, however, that all of these and similar terms are to be associated with the appropriate physical quantities and are merely convenient labels applied to these quantities. Unless specifically stated otherwise, or as is apparent from the discussion, terms such as processing or computing or calculating or determining of displaying or the like, refer to the action and processes of a computer system, or similar electronic computing device, that manipulates and transforms data represented as physical, electronic quantities within the computer system’s registers and memories into other data similarly represented as physical quantities within the computer system memories or registers or other such information storage, transmission or display devices.

[0127] Note also that the software implemented aspects of the example implementations are typically encoded on some form of non-transitory program storage medium or implemented over some type of transmission medium. The program storage medium may be magnetic (e.g., a floppy disk or a hard drive) or optical (e.g., a compact disk read only memory, or CD ROM), and may be read only or random access. Similarly, the transmission medium may be twisted wire pairs, coaxial cable, optical fiber, or some other suitable transmission medium known to the art. The example implementations are not limited by these aspects of any given implementation.

[0128] Lastly, it should also be noted that whilst the accompanying claims set out particular combinations of features described herein, the scope of the presentdisclosure is not limited to the particular combinations hereafter claimed, but instead extends to encompass any combination of features or implementations herein disclosed irrespective of whether or not that particular combination has been specifically enumerated in the accompanying claims at this time.

Claims

WHAT IS CLAIMED IS:

1. A method comprising: receiving a plurality of video streams; generating a plurality of frames of the plurality of video streams; generating a plurality of region of interest maps associated with the plurality of frames; and compressing the plurality of frames as a plurality of compressed frames including quantizing the plurality of frames using a quantization parameter associated with a corresponding one of the plurality of region of interest maps.

2. The method of claim 1 further comprising streaming the plurality of compressed frames.

3. The method of claim 1 or claim 2, wherein the plurality of video streams have different view perspectives of a scene.

4. The method of any of claim 1 to claim 3, wherein the generating of the plurality of region of interest maps includes using a model trained to indicate a redundancy between view perspectives associated with the plurality of frames.

5. The method of any of claim 1 to claim 4, wherein the generating of the plurality of region of interest maps includes using a model trained to indicate a foreground and a background associated with the plurality of frames.

6. The method of any of claim 1 to claim 5, wherein the generating of the plurality of region of interest maps includes using a model trained to indicate regions of the plurality of frames that include data necessary information for image synthesis.

7. The method of any of claim 1 to claim 6, wherein the generating of the plurality of region of interest maps includes using a model trained to indicate regions of the plurality of frames that that are not used for image synthesis.

8. The method of any of claim 1 to claim 7, wherein the plurality of region of interest maps are generated as a heat map where a value of a pixel of the heat map represents an interest level.

9. The method of claim 8, wherein the quantization parameter corresponds to the interest level.

10. The method of claim 8 further comprising determining the quantization parameter based on a look-up table and the value of the pixel.

11. The method of any of claim 1 to claim 10, wherein the generating of the plurality of region of interest maps includes using a model trained using a differential codec proxy, a quantization proxy, and a rate proxy.

12. The method of any of claim 1 to claim 11, wherein the generating of the plurality of region of interest maps includes using a model trained using a loss based on a bitrate loss and a distortion loss.

13. A non-transitory computer-readable storage medium comprising instructions stored thereon that, when executed by at least one processor, are configured to cause a computing system to: receive a plurality of video streams; generate a plurality of frames of the plurality of video streams; generate a plurality of region of interest maps associated with the plurality of frames; and compress the plurality of frames as a plurality of compressed frames including quantizing the plurality of frames using a quantization parameter associated with a corresponding one of the plurality of region of interest maps.

14. The non-transitory computer-readable storage medium of claim 13, wherein the instructions are further configured to cause the computing system to stream the plurality of compressed frames.

15. The non-transitory computer-readable storage medium of claim 13 or claim14, wherein the plurality of video streams have different view perspectives of a scene.

16. The non-transitory computer-readable storage medium of any of claim 13 to claim 15, wherein the generating of the plurality of region of interest maps includes using a model trained to indicate a redundancy between view perspectives associated with the plurality of frames.

17. The non-transitory computer-readable storage medium of any of claim 13 to claim 16, wherein the generating of the plurality of region of interest maps includes using a model trained to indicate a foreground and a background associated with the plurality of frames.

18. The non-transitory computer-readable storage medium of any of claim 13 to claim 17, wherein the generating of the plurality of region of interest maps includes using a model trained to indicate regions of the plurality of frames that include data necessary information for image synthesis.

19. The non-transitory computer-readable storage medium of any of claim 13 to claim 18, wherein the generating of the plurality of region of interest maps includes using a model trained to indicate regions of the plurality of frames that that are not used for image synthesis.

20. The non-transitory computer-readable storage medium of any of claim 13 to claim 19, wherein the plurality of region of interest maps are generated as a heat map where a value of a pixel of the heat map represents an interest level.

21. The non-transitory computer-readable storage medium of claim 20, wherein the quantization parameter corresponds to the interest level.

22. The non-transitory computer-readable storage medium of claim 20, wherein the instructions are further configured to cause the computing system to determine the quantization parameter based on a look-up table and the value of the pixel.

23. The non-transitory computer-readable storage medium of any of claim 13 to claim 22, wherein the generating of the plurality of region of interest maps includes using a model trained using a differential codec proxy, a quantization proxy, and a rate proxy.

24. The non-transitory computer-readable storage medium of any of claim 13 to claim 23, wherein the generating of the plurality of region of interest maps includes using a model trained using a loss based on a bitrate loss and a distortion loss.

25. An apparatus comprising at least one processor and at least one memory including computer program code, the at least one memory and the computer program code configured to, with the at least one processor, cause the apparatus at least to perform the method of any of claims 1-12.

26. An apparatus comprising means for performing the method of any of claims 1-12.

Citation Information

Patent Citations

  • Method and apparatus for encoding and decoding one or more views of a scene

    JP2023542860A

  • Method and system for optimizing image and video compression for machine vision

    US11533484B1

  • System and method for depth based adaptive streaming of video information

    US20140321561A1

  • Generating heat maps using dynamic vision sensor events

    US20190007678A1

  • Region-of-interest aware video coding

    US9167255B2