Variable-bandwidth video communication using a generative model
Generative codec models in video communication devices adaptively manage bandwidth and compute resources to maintain high-quality 3D communication by dynamically encoding and decoding content frames, addressing variable network constraints.
Patent Information
- Application Number
- PCT/US2024/035615
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-06-26
- Publication Date
- 2026-01-02
AI Technical Summary
Conventional video communication technologies face challenges in managing variable network bandwidth, leading to inconsistent quality of service, particularly in immersive 3D communication, where large data volumes require finite and dynamically varying network resources.
Implementing generative codec models with customizable layers in video communication devices to adaptively encode and decode content frames based on available bandwidth, balancing compute and network resources to maintain quality.
Ensures consistent high-quality 3D communication by dynamically managing bandwidth and compute resources, reducing data transmission while preserving visual and audio detail, thus optimizing resource usage and maintaining immersive experiences.
Smart Images

Figure US2024035615_02012026_PF_FP_ABST
Abstract
Description
VARIABLE-BANDWIDTH VIDEOCOMMUNICATION USING A GENERATIVE MODELBACKGROUND
[0001] Remote communication technologies have long provided ways for people separated by distance to communicate with one another in real time. For example, the telephone provided a form of real-time audio communication between remotely separated users for decades before other more sophisticated technologies (e g., cellular technologies, voice-over-Intemet-protocol (VoIP) technologies, etc.) were developed to let people communicate with more flexibility, lower cost, higher quality of service, and so forth. More recently, video communication technologies have become widely used to allow people to not only hear, but also see, one another in real-time over long distances.SUMMARY
[0002] Video communication generally, and emerging three-dimensional (3D) communication technologies more particularly, tend to involve the exchange of large amounts of data between video communication devices. For example, content frames representing a first scene may be transmitted from a first video communication device at the first scene to a second video communication device at a second scene where the content frames can be presented. While vary ing availability of resources (e.g., network bandwidth, etc.) may be managed by dropping content frames or employing conventional compression and encoding schemes (e.g., adaptive bitrate streaming, etc.), these resource management approaches may result in conspicuously low and / or variable quality of service during the communication session. Generative codec models described herein may therefore be used to manage a tradeoff between compute resources (e.g., processing and memory resources, which tend to be fairly abundant in modem video communication devices) and network bandwidth resources (which tend to be more limited). More particularly, a generative encoder on a transmitting side of communication link may, based on a current bandwidth availability, select a codec profile that processes a content frame through a certain number of layers (e.g., convolutional layers of a convolutional neural network implementing the generative codec model). This reduced amount of data may be transmitted across the network and reconstructed on a receiving side by a generative decoder that uses the same codec profile toreproduce the content frame with a minimal loss of quality.
[0003] To this end, one implementation described herein involves a method that may be performed by a first video communication device at a first scene. This example method may include, for instance: 1) generating a content frame for a visual representation of the first scene; 2) selecting, based on a bandwidth available for the first video communication device to exchange data with a second video communication device at a second scene, a codec profile to be used by a generative codec model, the codec profile defining a transmission layer from a series of layers included in the generative codec model; 3) encoding the content frame, by the generative codec model using the codec profile, in a pipeline extending from an input of an initial layer of the series of layers to an output of the transmission layer; and 4) transmitting, to the second video communication device, the encoded content frame with metadata indicating the transmission layer.
[0004] While this implementation describes the encoding / transmitting of communication data from the perspective of the first video communication device, it will be understood that the second video communication device may perform similar functions at the same time. For example, the second video communication device may also generate content frames, select an appropriate codec profile to be used by its generative codec model, encode the content frames using the codec profile, and transmit the encoded content frames to the first video communication device.
[0005] Moreover, both the first and second video communication devices, as they are encoding and transmitting content frames in accordance with methods described above, may receive encoded content frames sent by the other device and decode and present these content frames to implement the communication. For example, another implementation described herein involves a method performed by the second video communication device at the second scene in response to the first video communication device transmitting the encoded content frame with the metadata as described above. Specifically, this example method may include, for instance: 1) receiving, from the first video communication device, the encoded content frame for the visual representation of the first scene, the encoded content frame indicating a transmission layer of a series of layers included in a generative codec model; 2) selecting, based on the transmission layer indicated in the encoded content frame, a codec profile to be used by the generative codec model; 3) reconstructing, by the generative codec model using the codec profile, a content frame by decoding the encoded content frame in a pipeline extending from an input of the transmission layer to an output of a final layer of the series of layers; and 4) presenting, by the second video communication device, the content frame.
[0006] Other example implementations described herein involve the video communication devices themselves. For example, a first video communication device may include: a memory storing instructions and one or more processors communicatively coupled to the memory and configured to execute the instructions to perform a process. For example, the process may include the same or similar operations as described for the example encoder method above, such as: 1) generating a content frame for a visual representation of a first scene; 2) selecting, based on a bandwidth available for the video communication device to exchange data with an additional video communication device at a second scene, a codec profile to be used by a generative codec model, the codec profile defining a transmission layer from a series of layers included in the generative codec model; 3) encoding the content frame, by the generative codec model using the codec profile, in a pipeline extending from an input of an initial layer of the series of layers to an output of the transmission layer; and 4) transmitting, to the additional video communication device, the encoded content frame with metadata indicating the transmission layer.
[0007] Another example implementation described herein is for the second video communication device. For instance, like the first video communication device, the second video communication device may include a memory storing instructions and one or more processors communicatively coupled to the memory and configured to execute the instructions to perform a process. In this example, the process may include the same or similar operations as described above for the decoder-side method, such as: 1) receiving, at a second scene from an additional video communication device at a first scene, an encoded content frame for a visual representation of the first scene, the encoded content frame indicating a transmission layer of a series of layers included in a generative codec model; 2) selecting, based on the transmission layer indicated in the encoded content frame, a codec profile to be used by the generative codec model; 3) reconstructing, by the generative codec model using the codec profile, a content frame by decoding the encoded content frame in a pipeline extending from an input of the transmission layer to an output of a final layer of the series of layers; and 4) presenting, by the second video communication device, the content frame.
[0008] In still other implementations, any of these methods and / or processes may be embodied on non-transitory computer-readable media. For example, a non-transitory computer-readable medium may store instructions that, when executed (e.g., by one or more processors of a video communication device), cause the one or more processors to perform a process such as the encoder-side method set forth above. Similarly, the same or another non-transitory computer-readable medium may store instructions that, when executed, cause the same or another processor of a video communication device (e.g., the same video communication device or an additional video communication device) to perform a process such as the decoder-side method set forth above.
[0009] Various additional operations may be added to these processes and methods as may serve a particular implementation, examples of which will be described in more detail below. Additionally, it will be understood that each of the processes and operations described as being performed by different types of implementations in the examples above (e.g., the methods, the video communication devices, the non-transitory computer readable media, etc.) may additionally or alternatively be performed by other types of implementations as well. For example, a process described above as being embodied by a computer readable medium could be performed as a method and could be performed by a processor of a video communication device. Similarly, a method set forth above could be encoded in instructions stored by a computer readable medium or otherwise stored within the memory of a video communication device, and so forth.
[0010] The details of these and other implementations are set forth in the accompanying drawings and the description below. Other features will also be made apparent from the following description, drawings, and claims.BRIEF DESCRIPTION OF THE DRAWINGS
[0011] FIG. 1 shows illustrative aspects of an example implementation of variablebandwidth video communication using a generative codec model in accordance with principles described herein.
[0012] FIG. 2 shows an illustrative video communication device implementing a generative codec model in accordance with principles described herein.
[0013] FIGS. 3A-3B show illustrative encoder-side methods for variable-bandwidth video communication using a generative codec model in accordance with principles described herein.
[0014] FIGS. 4A-4B show illustrative decoder-side methods for variable-bandwidth video communication using a generative codec model in accordance with principles described herein.
[0015] FIG. 5A shows illustrative aspects of an example generative encoder included within a generative codec model in accordance with principles described herein.
[0016] FIG. 5B shows illustrative aspects of an example generative decoder includedwithin a generative codec model in accordance with principles described herein.
[0017] FIG. 6 shows illustrative aspects of how a codec profile may be selected by a generative encoder in accordance with principles described herein.
[0018] FIG. 7 shows illustrative aspects of how a codec profile may be selected by a generative decoder in accordance with principles described herein.
[0019] FIG. 8 shows illustrative scenarios in which a generative codec model uses different codec profiles that have been selected in accordance with principles described herein.
[0020] FIG. 9 shows an illustrative computing system that may be used to implement various devices and / or systems described herein.DETAILED DESCRIPTION
[0021] Implementations described herein may be applicable to any form of video communication, but may be especially useful for three-dimensional (3D) video communication due to large volumes of data that are typically exchanged for such communication. 3D video communication technologies continue the trend, described above, of providing users with ever more immersive and realistic ways to communicate over long distances. For example, 3D communication technologies (sometimes referred to by other terms such as teleportation, telepresence, holoportation, spatial conferencing, etc.) may help create more immersive and realistic experiences for remote communication by using 3D capture and projection techniques to create, in some examples, life-sized, 3D images of people and scenes in real-time. While significant resources (e.g., processing resources, communication bandwidth, etc.) may be required to implement it, 3D communication may allow participants in different locations to see and interact with each other in a manner that very much feels like in-person interactions they might have if they were physically present in the same space.
[0022] As used herein, 3D communication refers to communication between users in different places (e.g., different buildings, different cities, etc.) that utilizes 3D representations of the participating users rather than flat representations such as may be provided by a standard video call (e.g., with 2D video, mono or stereo sound, etc.). For example, in a 3D teleconferencing application, 3D communication technology may be used to create, between two 3D communication devices, a 3D teleconferencing session in which parties (e.g., single- user or multi-user parties) in different locations may see and hear life-sized, fully 3D representations of one another so that both parties can interact in a way that is experienced asbeing physically together at a same location.
[0023] To provide the added layer of immersion and realism beyond conventional video conferencing, 3D communication links may involve the transfer and processing of considerably more data than a conventional video link. As an example, subjects present in a scene (e.g., people, objects in the room, etc.) may be captured by a 3D capture mechanism within a 3D communication device. For instance, a plurality of cameras, depth scanning equipment, directional microphones, and so forth, may be used to form a 3D representation of the scene that can be presented on a 3D display screen such as a light field display or another suitable 3D display configured to present such a 3D representation (e.g., an autostereoscopic display, a holographic screen, etc.).
[0024] A given 3D communication device may be appropriately equipped with sufficient processing capability to handle the volume of data that is anticipated from whatever cameras, microphones, etc., the system includes. The transmission and exchange of the resulting data structures (e.g., content frames including 3D representations constructed from the data captured by all the cameras and microphones), however, may present a significant technical problem or challenge. For example, communication networks that may be used to carry data communications to and from a video communication device have a finite amount of bandwidth that may vary from moment to moment (e.g., based on usage of the network by other devices and other factors that are not within the control or influence of the video communication devices) and that may not be readily expandable. If a video communication device is used in a home or workspace, for example, the amount of data throughout available to it may be limited by the internet speed provided to the site, the Wi-Fi bandwidth on a wireless local area network (WLAN), and other such restraints that are outside the control of video communication device, regardless of how it is designed. If the network bandwidth available at a particular moment is insufficient to carry the large amounts of data produced by the video communication device, negative effects may be experienced such as delays in the communication (e.g., lagging or latency issues), unreliable communication (where the picture and / or sound periodically drops out as data buffers or loads), and so forth. Moreover, this technical problem may not only be limited to the video communication application but may also affect other systems and devices relying on the limited bandwidth as the video communication application dominates the available resources.
[0025] Implementations described herein present at least one technical solution to the technical problem of providing video communication using finite resources (using limited network bandwidth, in particular). Such solutions may be particular useful as it may not bepossible or practical to address the challenge by significantly increasing certain ty pes of available resources. For example, the Wi-Fi and / or network bandwidth available for transferring communication data at a particular site may be fixed to a certain extent and to increase the bandwidth to facilitate 3D video communication may not be possible or desirable. Moreover, conventional approaches to addressing limited communication bandwidth (e.g., reducing the frame rate or otherwise dropping frames, downshifting to more heavily compressed and lower quality versions of the frames such as in an adaptive bitrate streaming (ABR) scheme, etc.) may lead to noticeable quality7degradation or other undesirable outcomes. Accordingly, implementations described herein relate to variablebandwidth video communication using generative codec models. As detailed herein, these models may allow video communication devices to reduce the total amount of data being transmitted and received, even while allowing the content that is ultimately presented (e.g., 3D video content, spatial 3D audio, etc.) to be perceived as having the same or a similar level of detail and quality7as it would have if there were unlimited network bandwidth to exchange all the communication data being produced.
[0026] As will be described and illustrated in detail below, generative codec models may be implemented as custom U-nets with common layer loss and may provide the functionality7and solutions described above by offering a novel framework to adaptively interpolate visual and audio content (e.g., 3D content frames, etc.) based on a generative scheme of custom encoder-decoder networks. For example, a corresponding series of layers (e.g., convolutional deep neural network layers) within both the encoder model (implemented in one video communication device) and the decoder model (implemented in the other video communication device being communicated with) may use processing resources and prior training to generatively and progressively encode content frames to be represented by vectors with fewer and fewer dimensions to satisfy the amount of bandwidth that is presently available for data transmission. When more bandwidth is available, an earlier stopping point in this series of layers may be selected and used as a transmission layer (with further layers going unused for that data). On the other hand, when less bandwidth is available, a later transmission layer in the pipeline may be selected as the transmission layer, such that more or all of the available layers of the encoder and the decoder are used to minimize the dimensionality7of the representation sent through the network.
[0027] Ultimately, the technical effect of variable-bandwidth video communication using generative codec models described herein is that the system may dynamically select levels of extrapolation and communication bandwidth used based on resources that areavailable at any given time. In other words, generative codec models may allow for opportunistic selection of the size of the encoding to upscale and downscale with similar end results, but with more or less data processing and transmission being performed. Indeed, based on current resource availability, video communication devices may be configured to effectively manage a tradeoff betw een compute resources and network bandwidth resources as may best serve the overall system at any given moment.
[0028] For example, when the main bottleneck is network bandwidth, the system may dynamically respond to that by using more layers (and processing resources) while reducing bandwidth usage and still maintaining a threshold level of quality. Conversely, when more network bandwidth is available, or if compute resources such as memory and processing become the bottleneck, the system may dynamically respond by using fewer layers (and processing resources) and increasing the amount of data transmitted to maintain the quality of service. In either case, users may experience a consistently high quality of senice throughout their 3D communication session while resources are conserved and used efficiently. In particular, the total amount of data being transmitted from a given video communication device may be at a level that can be readily handled by existing infrastructure while still leaving resources for other applications and use cases. At the same time, a reliable and consistent communication experience may be enjoyed without noticeable lag or latency, without the picture or sound dropping out. and otherwise without issues that would detract from the immersive communication experience that is being provided.
[0029] Various implementations will now be described in more detail with reference to the figures. It will be understood that particular implementations described below are provided as non-limiting examples and may be applied in various situations. Additionally, it will be understood that other implementations not explicitly described herein may also fall within the scope of the claims set forth below. Systems and methods described herein for variable-bandwidth video communication using a generative codec model may result in any or all of the technical effects mentioned above, as well as various additional effects and benefits that will be described and / or made apparent below.
[0030] FIG. 1 shows illustrative aspects of an example implementation 100 of variable-bandwidth video communication using a generative codec model in accordance with principles described herein. As shown, implementation 100 depicts a communication session between a first video communication device 102-1 and a second video communication device 102-2. Video communication device 102-1 is shown to have a first display screen 104-1. to be located at a first scene 106-1, and to be used by a first user 108-1. Similarly, videocommunication device 102-2 is shown to have a second display screen 104-2, to be located at a second scene 106-2, and to be used by a second user 108-2.
[0031] During the communication session, user 108-1 may be presented, on display screen 104-1, a view of scene 106-2 and user 108-2 within that scene. In some examples, display screen 104-1 may be implemented as a 3D display (e.g., an autostereoscopic display such as a light field display, a holographic display, etc.), such that the representation of user 108-2 viewed by user 108-1 is a 3D representation (e.g., a life-sized representation that helps immersively simulate an in-person interaction between the users). Similarly, as further shown in FIG. 1, user 108-2 may be presented, on display screen 104-2, a view of scene 106-1 and user 108-1 within that scene. Like display screen 104-1, display screen 104-2 may be implemented as a 3D display, such that the representation of user 108-1 viewed by user 108-2 is a 3D representation (e.g., a similar life-sized representation to help simulate the in-person interaction).
[0032] The communication session betw een video communication devices 102-1 and 102-2 is shown to be enabled by a data exchange 110 between the devices that is carried out over a network 1 12. For example, network 112 may represent any suitable local, wide area, public, private, and / or other communication networks as may be included on a data path between the two video communication devices. Such netw orks included within network 112 could include, for example, wireless (e.g., Wi-Fi) networks, local area netw orks (LANs), wide area networks (WANs), carrier networks (e.g.. cellular networks), private or public networks, the internet, and / or any other networks as may be used to cany' data communications between the video communication devices. Whatever networks or netw ork elements may be included within network 112 between video communication devices 102-1 and 102-2, network 112 will be understood to have a finite and dynamically varying amount of available bandwidth for data exchange 110 between the devices. For example, if video communication devices 102-1 and / or 102-2 are used in homes or w orkspaces, the amount of network bandwidth (e.g., data throughout) available on network 112 may be limited by the Internet speed provided to these sites, the Wi-Fi bandwidth on a wireless local area network (WLAN) at the sites, and other such restraints that may be outside the control of the devices. This available bandwidth may also vary dynamically as other devices (not shown in FIG. 1) share the network bandwidth and as other external conditions affect the rate of data exchange 110 that netw ork 112 can support.
[0033] Accordingly, FIG. 1 shows that data exchange 110 between the video communication devices 102-1 and 102-2 may be facilitated by use of respective generativecodec models implemented by each video communication device. Specifically, as shown on either side of the dashed line separating scene 106-1 from scene 106-2 in the figure, video communication device 102-1 may implement a generative codec model 1 14-1 and video communication device 102-2 may implement a generative codec model 114-2, each of which supports a plurality of codec profiles 116. While generative codec models 114-1 and 114-2 are illustrated as separate models in FIG. 1, it will be understood that these models may be related and configured (e.g.. trained) to work together to properly encode and decode communication data for data exchange 110 (e.g., content frames associated with the communication session). Indeed, to support bidirectional, full-duplex communication between the video communication devices, arrows in FIG. 1 show' both data that is being encoded and transmitted, as well as data that is being received and decoded, by both video communication devices on both sides of the communication link.
[0034] A stream of content frames 118-1 generated for a visual representation of scene 106-1 (illustrated as individual boxes) is shown to be entering generative codec model 114-1 to then be converted into encoded content frames 120 and transmitted over network 112 (e.g.. as part of data exchange 110). These encoded content frames 120 are received by generative codec model 114-2 and reconstructed (e.g., decoded) into content frames 118-2 that can then be presented by video communication devices 102-2 to present a representation of scene 106-1 (including user 108-1) to user 108-2 on display screen 104-2. At the same time, other content frames 118-2 generated for a visual representation of scene 106-2 are shown to be entering generative codec model 1 14-2 to then be converted into encoded content frames 120 and transmitted in the other direction over network 112 (e.g., also as part of data exchange 110). These encoded content frames 120 are received by generative codec model 114-1 and reconstructed (e.g., decoded) into content frames 118-1 that can be presented by video communication devices 102-1 to present a representation of scene 106-2 (including user 108-2) to user 108-1 on display screen 104-1. Accordingly, while various examples described and illustrated herein focus on the generative encoder of one generative codec model and the corresponding generative decoder of the other generative codec model, it will be understood that a full generative codec model such as generative codec model 114-1 or generative codec model 1 14-2 may include both a generative encoder and a generative decoder to support the simultaneous tw o-w ay communication shown in FIG. 1.
[0035] As described in more detail below, various codec profiles 116 supported by generative codec models 114-1 and 114-2 may allow for flexibility in how much network bandwidth of network 112 may be utilized at any given time. For example, as describedabove, different profiles may be selected to dynamically and efficiently manage a tradeoff between compute resources (e.g., processing and / or memory resources, etc.) that the generative codec models consume and network resources of network 112 that data exchange 110 uses. For example, one codec profile 116 (which may be used for a given content frame both on the encoder side and the decoder side) may be configured to consume more compute resources while requiring less network bandwidth, while another codec profile 116 may be configured to consume fewer compute resources while requiring more network bandwidth. As will be described, selection between these and other codec profiles 116 that balance the tradeoff in various w ays may be made dynamically by video communication devices 102-1 and 102-2 based on what resources (e.g., bandwidth and / or compute resources) happen to be available at a given time.
[0036] Whichever of a set of supported codec profiles 116 may be selected for the data transfer at a given time, it will be understood that the usage of these different profiles to manage the tradeoff betw een compute and bandwidth resources is distinguishable in several respects from conventional bandwidth management techniques (e.g., dropping frames, changing frame rates, adaptive bitrate streaming, etc.). As will be described in more detail below, generative codec models 114-1 and 114-2 may be trained for communication sessions in an end-domain specific way. In other words, the training may be rich in semantics to help the neural netw orks master the end-to-end behavior of the types of content frames that are consistently processed for the 3D video communication application or use case. While a conventional image encoder may, for example, analyze pixel groupings to make correlations and perform transforms to compress the image data in a generic manner configured to be applied to all types of images, generative codec models 114-1 and 114-2 may be trained, at each of a plurality of layers, to generatively produce low-dimensional embeddings representing the input from the previous layer and to effectively serve content frames of the type produced by the video communication devices. This may include various embedded assumptions that are trained into the model, such as that content frames will depict relatively static scenes from frame to frame, that dynamic human subjects are likely to be in depicted in the middle of each frame, that there is more important texture to capture in the middle of the frame (where the human subject is) and less texture around the edges, and so forth.
[0037] As a result of this training, generative codec models may be configured to produce progressively low-dimensional embeddings to represent content frames depending on how many layers of processing a given frame goes through prior to reaching a transmission layer that is selected to be the final layer before the frame is transmitted over thenetwork. When appropriate, many layers of processing could produce a vector with a relatively low number of dimensions (e.g., less than 1000 values) to be transferred over the network. While a profile with so many layers may require a significant amount of compute resources, all the visual semantics of the content frame may be preserved at a very high level of quality given the immense reduction of data that is actually transferred through the bottleneck of network 112. On the other hand, when more bandwidth is available, a profile with fewer layers of generative processing may alleviate the compute resources needed by the models and may achieve the desired quality by transferring a larger amount of data.
[0038] FIG. 2 shows an implementation 200 of an illustrative video communication device 102 implementing a generative codec model 114 in accordance with principles described herein. For example, implementation 200 may represent either or both of video communication devices 102-1 and 102-2 described above in relation to FIG. 1. Video communication device 102 is shown to illustrate certain elements that may be present in a given implementation of a video communication device in accordance with principles described herein. It will be understood, however, that fewer or additional elements may be included in other implementations of video communication device 102 and that implementation 200 is offered only as an example for illustrative purposes.
[0039] As shown, implementation 200 of video communication device 102 includes a display screen 104 that may present visual data for a communication session. For example, as described and illustrated above, display screen 104 may be used for presenting content frames by displaying representations (e.g., 2D representations, 3D representations, etc.) that the content frames depict. In some examples, display screen 104 may be implemented as a 3D display (e.g., an autostereoscopic display, etc.) configured to display a 3D representation of a subject present at a scene. In this way. a user such as user 108-2 may be presented with a 3D representation of another user such as user 108-1 without needing to wear 3D glasses or the like.
[0040] Video communication device 102 is further shown to include an array of cameras that may be used to capture images from which such content frames are produced. For example, video communication device 102 may be configured to generate a content frame for a visual representation of a scene by performing operations including, for instance: 1) capturing, from a plurality of vantage points by the plurality of cameras 202, a plurality7of images of the scene (e.g., scene 106-1); and 2) constructing, based on the plurality of images, a 3D representation of a subject present at the scene (e.g.. user 108-1). The visual representation of the scene included in the content frame may therefore include the 3Drepresentation of this subject. The capture of the images and the construction of the 3D representation may be performed in any manner as may serve a particular implementation (e.g., using known 3D scanning and modeling techniques and technologies, etc.).
[0041] Video communication device 102 is further shown to include an array of loudspeakers 204 and an array of microphones 206. While sound capture and reproduction are not explicitly illustrated in FIG. 1 and is not a focus of the present disclosure, it will be understood that 3D sound (e.g., spatial audio, etc.) may be captured using multiple microphones and reproduced using multiple loudspeakers in any manner as may sen e a particular implementation. In some examples, content frames described herein may include not only 3D visual representations of the scene but also 3D audio associated with the scene. In other examples, audio frames may be handled in a separate communication pipeline from visual content frames described herein.
[0042] Video communication device 102 is also shown to include various processors including one or more processors 208 that may represent general purpose processors (e.g., central processing units (CPUs), microprocessors, etc.) or more special purpose processors (e.g., graphics processing units (GPUs), field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), etc.). In the example of implementation 200, video communication device 102 is also shown to include one or more machine learning acceleration processors 210 that will be understood to be specialized processors configured to facilitate computation performed by generative codec model 114. For example, machine learning acceleration processors 210 may be specialized for machine learning workloads that involving large amounts of data and repetitive mathematical operations. To this end, machine learning acceleration processors 210 may have many cores (e g., thousands of cores in some examples, or at least many more cores than general purpose processors may have, etc.) that are optimized for specific machine learning calculations (e.g., matrix multiplication operations, etc.). In this way, machine learning acceleration processors 210 may support and / or provide high throughput for specialized tasks involving large machine learning matrices. Machine learning acceleration processors 210 may also have limited instruction sets compared to general purpose processors, since the instruction sets may focus specifically on instructions heavily used in machine learning algorithms (e.g., operations involving processing massive datasets, accelerating computations in machine learning tasks like image recognition, etc.).
[0043] Video communication device 102 is also shown to include a memory 212 that may store instructions and may be communicatively coupled to the various processors 208and 210 so that the processors may execute the instructions. Certain instructions and other data stored in memory 212 may embody a generative codec model 114 with any of the characteristics and / or functions described herein (e.g., characteristics and / or functions described above for generative codec models 114-1 and 114-2, etc.). As shown, generative codec model 114 may include both: 1) a generative encoder 214 for encoding content frames and preparing encoded content frames for transmission to another video communication device, and 2) a generative decoder 216 for decoding encoded content frames that have been received from the other video communication device. Other instructions stored in memory 212 may be configured for execution by processors 208 and / or 210 to perform any of various processes 218. For example, such processes may correspond to methods that will now be described in relation to FIGS. 3A-3B and FIGS. 4A-4B.
[0044] FIGS. 3A and 3B show illustrative encoder-side methods (i.e. , a method 300 in FIG. 3 A and a method 310 in FIG. 3B) for variable-bandwidth video communication using a generative codec model in accordance with principles described herein. FIGS. 4A and 4B then show illustrative decoder-side methods (i.e., a method 400 in FIG. 4A and a method 410 in FIG. 4B) for variable-bandwidth video communication using the generative codec model in accordance with principles described herein. While, as described and illustrated above in relation to FIG. 1, both video communication devices 102-1 and 102-2 may be capable of the same encoding / decoding functionality, specific examples described herein, including for methods 300, 310, 400, and 410. will assume that video communication device 102-1 is encoding and transmitting communication data while video communication device 102-2 is receiving and decoding that data. Accordingly, methods 300 and 310 will be described as being performed by video communication device 102-1, though it will be understood that the same methods could also be performed by video communication device 102-2 and / or other implementations of video communication device 102 described herein (e.g., in some examples, being performed simultaneously so as to implement full-duplex video communication). Similarly, methods 400 and 410 will be described as being performed byvideo communication device 102-2, though it will be understood that the same methods could also be performed by video communication device 102-1 and / or other implementations of video communication device 102 described herein.
[0045] While FIGS. 3A-3B and 4A-4B show illustrative operations according to specific implementations, it will be understood that other implementations of these methods may omit, add to, reorder, and / or modify any of operations that are explicitly represented in FIGS. 3A-3B and 4A-4B. Additionally, while operations shown in these figures areillustrated with arrows suggestive of a sequential order of operation, it will be understood that some or all of the operations of methods 300, 310, 400, and 410 may be performed concurrently (e.g., in parallel) with one another. Each of the operations of these methods will now be described in more detail as the operations may be performed by the respective implementations of video communication device 102 noted above (i. e. , video communication device 102-1 for methods 300 and 310, video communication device 102-2 for methods 400 and 410).
[0046] At operation 302, video communication device 102-1 may generate, at a first scene (e.g., scene 106-1), a content frame for a visual representation of the first scene. For example, as has been described, video communication device 102-1 may use a plurality of cameras (e.g., cameras 202) positioned at different vantage points to capture image and depth data for the scene, including for any human subjects that are present at the scene (e.g., user 108-1). In some examples, audio (e.g., stereo or spatial audio) may be similarly captured using a plurality7of microphones (e.g., directional microphones such as microphones 206) placed at the scene. Based on such image, depth, and / or audio data, video communication device 102-1 may generate a 3D representation of the scene and / or of the human subjects at the scene, in particular. The content frame may include data associated with that 3D representation for one or more frame times. For instance, the content frame may include visual (e.g., image) data, depth data, texture data, audio data, metadata, and / or any other information as may be useful to reproduce the visual representation of the first scene as captured, as well as corresponding audio in certain examples.
[0047] At operation 304, video communication device 102-1 may select a codec profile to be used by a generative codec model. For example, the selection at operation 304 may be performed based on a bandwidth available for video communication device 102-1 to exchange data with video communication device 102-2 at a second scene (e.g., scene 106-2). Moreover, other bases may also be accounted for in the selection process, such as an availability of compute resources such as processing and / or memory7resources. The codec profile selected to be used by the generative codec model may be any of the plurality of codec profiles 116 supported by generative codec model 114-1 described above. For example, the selected codec profile may manage a tradeoff between network bandwidth and compute resources by defining a particular layer (e.g., a convolutional neural network layer), from a series of layers included in generative codec model 114-1, as a transmission layer. As has been described, the transmission layer selected and defined by the codec profile may refer to a final layer in the series of layers that is used to process the content frame beforetransmitting the content from over the network using the available bandwidth.
[0048] Accordingly, at operation 306, video communication device 102-1 may encode the content frame. More particularly, the generative codec model 114-1 within video communication device 102-1 may, based on the codec profile selected at operation 304, encode the content frame in a pipeline that extends from an input of an initial layer of the series of layers to an output of the transmission layer. For example, if there were ten layers in the series of layers and layer 4 were selected as the transmission layer, the encoding of the content frame at operation 306 may be performed by a pipeline that begins with layer 1 (i.e., the initial layer in this example) and continues with layer 2, layer 3, and layer 4. While additional layers 5-10 may still be available to further encode the output of layer 4 to further reduce the dimensionality of the data that is to be sent over the network, the selection of layer 4 as the transmission layer for this example means that these further layers would be bypassed for this particular content frame (since the amount of data output by layer 4 may be appropriate given the current bandwidth available). If a different codec profile associated with a different transmission layer (e.g., layer 2. layer 10, etc.) were selected for a different example, the encoding at operation 306 would instead use a pipeline with a different number of layers (e.g., layers 1-2, layers 1-10, etc.). In some examples, transmission layer may be defined as the initial layer, such that the pipeline only includes a single layer (i.e., the initial / transmission layer). In this case, the pipeline would extends from the input of the initial layer to the output of the initial layer (which is, by definition, also the output of the transmission layer in this example).
[0049] At operation 308, video communication device 102-1 may transmit, to video communication device 102-2, the encoded content frame produced by operation 306. For instance, in the example where the codec profile defines layer 4 (of a ten-layer series) as the transmission layer, the encoded content frame may comprise the content frame as encoded by layers 1-4. To indicate how the encoded content frame is to be decoded upon receipt by video communication device 102-2, the encoded content frame transmitted at operation 308 may be transmitted to video communication device 102-2 with metadata indicating the transmission layer (e.g., layer 4 in this arbitrary example). For instance, the encoded content frame may be transmitted with a frame header or other metadata format that indicates which layer, of the series of layers, w as selected as the transmission layer so that the corresponding decoding layers (e.g., layers 4, 3, 2, and 1) may be used to reconstruct the content frame on the decode side. It will be understood that the metadata may indicate the transmission layer in any suitable way (e.g., using any format, etc.). In some examples, the transmission layer mayhave a unique relationship to the selected codec profile (i.e., such that there is a one-to-one correspondence between the supported codec profiles and the layers that may be used as transmission layers). In these examples, the metadata indicating the transmission layer could do so by indicating the codec profile selected at operation 304, since the transmission layer defined by that codec profile would be uniquely implied.
[0050] As has been mentioned, the bandwidth available for video communication device 102-1 to exchange data with video communication device 102-2 may not only be limited or finite, but it may also be variable and dynamic (i.e., changing over time). Depending on how often and by how much the bandwidth availability' changes, it may be useful for later content frames to be encoded and transmitted using different codec profiles than the codec profile selected at operation 304. For instance, if a codec profile defining the transmission layer as layer 4 was selected for one content frame (as per the arbitrary example described above), an additional codec profile defining the transmission layer as some other layer of the series other than layer 4 could be selected for a subsequent content frame when dynamic conditions have changed. As another arbitrary example using the illustrative series of ten layers, for instance, the bandwidth detected to be available could decrease by a certain amount and an additional codec profile could be selected that defines layer 7 as the transmission layer.
[0051] To illustrate more generally, FIG. 3B shows method 310, which is shown to pick up where method 300 left off. Specifically, sometime after encoding the content frame using the selected codec profile and transmitting this encoded content frame, video communication device 102-1 may determine that conditions such as the available network bandwidth have changed.
[0052] Accordingly, at operation 312, video communication device 102-1 may select an additional codec profile that defines an additional transmission layer from the series of layers. For example, if the bandwidth on which the selecting of the codec profile (at operation 304) is based is a first bandwidth at a first time during a communication session that includes the content frame and an additional content frame, this selection at operation 312 may be based on a second bandwidth at a second time during the communication session. The additional transmission layer defined by the additional codec profile may be different from the transmission layer. For instance, as mentioned in the example above, if the transmission layer defined by the earlier codec profile were layer 4, the additional transmission layer defined by the additional codec profile selected at operation 312 could be layer 7 (assuming that the second bandwidth was detected to be less than the first bandwidth).
[0053] At operation 314, video communication device 102-1 may encode the additional content frame in an analogous way as described above for the encoding of the content frame at operation 306. For example, the encoding may be performed by the generative codec model using the additional codec profile, such that the encoding is achieved in an additional pipeline extending from the input of the initial layer (e.g., layer 1) to an output of the additional transmission layer (e.g.. layer 7 in this example). As such, additional processing associated with layers 5, 6. and 7 may be performed for the additional content frame that were not performed for the content frame, thereby trading some compute resources (e.g., processing cycles, memon space, etc.) for an even lower dimensional encoded content frame (i.e., a smaller encoded content frame that can be communicated using less data and less network bandwidth).
[0054] At operation 316, video communication device 102-1 may then transmit the encoded additional content frame to video communication device 102-2. As with the transmission at operation 308, this transmission may also include metadata to indicate the additional transmission layer that will be useful for decoding and reconstructing the content frame when received by video communication device 102-2. Using these same principles, additional content frames may continue to be encoded and transmitted using dynamically changing codec profiles that manage the tradeoff between compute resources and network bandwidth in a manner configured to optimize efficiency and quality of service for the communication session.
[0055] Method 400 of FIG. 4A represents a decoder-side method corresponding to method 300 described above. For example, method 400 may be performed on the decoder side by video communication device 102-2 as video communication device 102-1 performs method 300 on the encoder side, as has been described.
[0056] At operation 402, video communication device 102-2 may receive (at the second scene) an encoded content frame transmitted (from the first scene) by video communication device 102-1. As has been described and illustrated, this encoded content frame (e.g., one of encoded content frames 120) may represent a visual representation of the first scene and may have been encoded, using a particular codec profile, by the generative encoder of the generative codec model in video communication device 102-1. For example, the encoded content frame received at operation 402 may be the encoded content frame described as being transmitted at operation 308 and, as such, may indicate (e.g., in the metadata of a header or the like) a transmission layer of a series of layers included in the generative codec model. As put forth in the running example above, for example, thisindicated transmission layer could be layer 4.
[0057] At operation 404. video communication device 102-2 may select a codec profile to be used by the generative codec model. For example, rather than assessing the available bandwidth and / or other present conditions, video communication device 102-2 may select the codec profile based on the transmission layer indicated in the encoded content frame that is received at operation 402 (e.g.. indicated within the metadata). In this way, even if conditions such as the bandwidth have changed since the encoded content frame was transmitted, the generative codec model may properly interpret and decode the encoded content frame in accordance with how the frame was encoded on the encoder side.
[0058] At operation 406. video communication device 102-2 may reconstruct a content frame. More particularly, the generative codec model implemented in video communication device 102-2 (e.g., generative codec model 114-2) may use the codec profile selected at operation 404 to decode the encoded content frame in a pipeline extending from an input of the transmission layer to an output of a final layer of the series of layers. In other words, the pipeline used to decode the encoded content frame may mirror the pipeline used to encode the frame, but in the reverse order. For example, if the transmission layer is layer 4, then the pipeline used by the generative decoder of video communication device 102-2 may begin by decoding the received frame using layer 4, then proceed to decode the output of that in layer 3, layer 2, and finally layer 1 (i.e.. the final layer of the series in this example). Similarly as described above, in some examples, the transmission layer and the final layer may be one and the same, such that the pipeline includes only a single layer and extends from the input of the transmission / final layer to the output of that same layer.
[0059] At operation 408, video communication device 102-2 may present the content frame that has been reconstructed. By nature of the corresponding layers used to generatively encode and then generatively decode the content frame, this presentation may be semantically identical or very similar to the content frame originally generated at operation 302. In other words, the content frame presented by video communication device 102-2 may be high quality and represent a very good representation of the first scene as captured by video communication device 102-1.
[0060] In a similar way as method 400 of FIG. 4A represented the decoder-side method corresponding to method 300, a method 410 in FIG. 4B will be understood to represent a decoder-side method corresponding to method 310. For example, method 410 may be performed on the decoder side by video communication device 102-2 as video communication device 102-1 performs method 310 on the encoder side, as has beendescribed.
[0061] Similar to method 310, method 410 is shown to pick up where method 400 left off. Specifically, sometime after decoding the encoded content frame received at operation 402 using the selected codec profile and presenting this content frame, conditions such as the available network bandwidth may change such that the generative encoder may begin using a different codec profile.
[0062] At operation 412. video communication device 102-2 may therefore select an additional codec profile corresponding to the new codec profile being used to encode and transmit the frames. For example, if the encoded content frame and an additional encoded content frame are both received as part of a communication session, the additional encoded content frame may indicate an additional transmission layer of the series of layers, the additional transmission layer different from the transmission layer. The selection of the additional codec profile at operation 412 may then be performed based on the additional transmission layer indicated in the additional encoded content frame. In other words, referring to the specific examples set forth above, the transmission layer may have been layer 4 and this may change to an additional transmission layer of layer 7. Accordingly, the codec profile associated with layer 7 may be selected as the additional codec profile at operation 412 (e.g., replacing the prior selection of the codec profile defining the transmission layer as layer 4).
[0063] At operation 414. video communication device 102-2 may reconstruct an additional content frame. Operation 414 may be performed by the generative codec model of video communication device 102-2 using the additional codec profile selected at operation 412. For instance, the generative codec model may decode the additional encoded content frame in an additional pipeline extending from the input of the additional transmission layer (e.g., layer 7 in the running example) to the output of the final layer of the series of layers (e.g., still layer 1). In this way, as described above in relation to operation 406, the multilayer encoding may be reversed by a pipeline with layers trained in all the same ways to ultimately produce the additional content frame as being very similar to the additional content frame generated at operation 312 of method 310.
[0064] At operation 416, video communication device 102-2 may present this additional content frame that has been reconstructed in the same way as the content frame was presented at operation 408. Again, by nature of the corresponding layers used to generatively encode and then generatively decode the additional content frame, this presentation may be semantically identical or very' similar to the content frame originallygenerated at operation 312. In other words, the content frame presented by video communication device 102-2 may again be high quality and represent a good representation of the first scene as captured by video communication device 102-1, though different amounts of bandwidth and processing resources may have been employed (using a different tradeoff ratio) to effectively communicate the content frame.
[0065] As was illustrated in implementation 200 of video communication device 102 in FIG. 2, a generative codec model such as generative codec model 114 may include both a generative encoder (e g., generative encoder 214) and a generative decoder (e.g., generative decoder 216). While the codec model may include both of these elements and correspondingly function to perform both encoding and decoding, however, it will be understood (as illustrated in FIG. 1) that content frames encoded by one generative codec model may be decoded by another generative codec model implemented in another video communication device. For example, given that both video communication devices 102-1 and 102-2 incorporate their own respective generative codec models 114-1 and 114-2, content frames encoded and transmitted by a generative encoder of generative codec model 114-1 may be received and decoded by a generative decoder of generative codec model 1 14-2. while content frames encoded and transmitted by a generative encoder of generative codec model 114-2 may be received and decoded by a generative decoder of generative codec model 114-1. To illustrate additional details for these respective encoding and decoding elements within the generative codec models 114. FIGS. 5A and 5B will now be described.
[0066] FIG. 5 A shows illustrative aspects of an example generative encoder 214 that may be included within a generative codec model such as generative codec model 114-1 (continuing the example convention described above that addresses content frames being communicated from video communication device 102-1 to video communication device 102- 2). FIG. 5B then shows illustrative aspects of a corresponding generative decoder 216 that may be included within a generative codec model such as generative codec model 114-2. As described above, it will be understood that a similar generative encoder may also be included in generative codec model 114-2 and that a similar generative decoder may also be included in generative codec model 114-1.
[0067] As shown in FIG. 5A, this implementation of generative encoder 214 may include an input 502 where content frames 118 (e.g., content frames 118-1 described in relation to FIG. 1) may be received, another input referred to as profile selection 504 that indicates which of a set of codec profiles 116 is to be used for the encoding (described in more detail below), and an output 506 where encoded content frames 120 can be output fortransmission to another video communication device. Additionally, generative encoder 214 is shown to include a series of layers 508 that includes an initial layer 508-1, and some number of other layers 508-2 through 508-N.
[0068] As has been described, each codec profile 116 that may be supported by a given generative codec model may define a transmission layer (a particular layer, from the series of layers 508, that is designated as the transmission layer for that codec profile). Depending on which codec profile is selected, then, the encoding of a given content frame 1 18 may be performed in a pipeline extending from an input of the initial layer (i.e., initial layer 508-1 in this example) to an output of the selected transmission layer (e.g., any of layers 508-1 through 508-N as may be associated with the selected codec profile). To illustrate the relationship between codec profiles 116 supported by this generative encoder 214 implementation and the series of layers 508, FIG. 5A depicts a variety of pipelines, each including a different pipeline of layers 508, for supported codec profiles 116. For example, a first codec profile 116-1 is shown to include a pipeline that is limited only to initial layer 508- 1. In this example, layer 508-1 is therefore not only the initial layer but also has been defined as the transmission layer for codec profile 116-1. A second codec profile 116-2 that defines layer 508-2 as the transmission layer is shown to include a pipeline that extends from initial layer 508-1 to layer 508-2. A third codec profile 116-3 that defines layer 508-3 as the transmission layer is similarly shown to include a pipeline that extends from initial layer 508- 1 to layer 508-3. Finally, in the limit, an Nth codec profile 116-N that defines layer 508-N as the transmission layer is shown to include a pipeline that extends from initial layer 508-1 to layer 508-N. In other words, the Nth codec profile 116-N will be understood to include the entire series of layers 508 so as to achieve the lowest bandwidth usage at the cost of the greatest amount of processing.
[0069] While not explicitly shown by the arrows in FIG. 5A, it will be understood that, for a given codec profile 116, layers 508 that come after the selected transmission layer in the series may not be used. For example, while codec profile 116-2 is selected, each content frame 118 may be processed by initial layer 508-1 and layer 508-2 and may then skip the remainder of the layers 508-3 through 508-N to be output directly as an encoded content frame 120 at output 506.
[0070] As has been mentioned, generative codec models such as generative codec model 114 may be specifically trained and configured for the particular application or use case for which they will be used (e.g., 3D video communication). As such, the generative codec model in this example may be trained using video communication data configured withsemantically relevant elements for a video communication application. These semantically relevant elements may eventually bestow the generative codec model with a thorough understanding of the types of content frames likely to be encountered. For example, these content frames may often include dynamic 3D models of human subjects that are more or less centered in the frame and include relatively high degrees of detail surrounded by more static backgrounds with lower degrees of detail, and so forth. By being trained with this specific understanding of the application specific semantics, each layer 508 may be configured to generate a representation of the content frame that has lower dimensionality than what the layer received while also being well configured for further processing by the next layer in the series or for transmission and decoding on the other side of the communication link.
[0071] Each layer (other than the initial layer) may be preceded by and followed by another layer with which it may be in communication. For example, the initial layer may provide a content frame representation to a second layer, which may receive that as input, perform further processing, and provide yet another representation (with an even lower dimensionality) to a third layer, and so forth. Accordingly, each layer may be specifically trained on the type of input it will receive from the layer before it in the generative codec model and may provide its output either to the next layer in the series or to a network interface (if it is selected as the transmission layer at a particular time) to be transmitted to the other communication device.
[0072] To this end, generative encoder 214 and the series of layers 508 may be implemented using any type of generative machine learning or artificial intelligence models as may serve a particular implementation. As one example, the generative codec model in which generative encoder 214 is included may be a convolutional neural network (CNN) and the series of layers 508 may be a series of convolutional layers within the neural network. In other examples, other types of models could be used instead or in combination with the CNN. For instance, a transform of a deep learning network, a support vector machine (SVM), a K- nearest neighbor (KNN)-based model, or another suitable model may be used. As will be described in more detail below, each layer 508 in the generative encoder 214 may be paired with a corresponding layer in the generative decoder 216 described below. During training, each of these encoder-decoder layer pairs may be configured to extrapolate visual / audio content and the overall network training loss may add the individual layer loss to ensure that each of the codec profiles 116 meets a particular loss threshold.
[0073] The N number of layers shown in FIGS. 5 A and 5B may represent any suitable number of layers as may serve a particular implementation. For example, certainimplementations may have a relatively small number of layers (e.g., 2 layers, 3 layers, 5 layers, etc.) so as to support a similarly small number of different codec profiles 116. Other implementations, however, may have a much larger number of layers in the series (e.g., 10 layers, 20 layers, 50 layers, 100 layers, etc.) to thereby support a larger number of codec profile options and a high degree of granularity with which to respond to different network conditions.
[0074] As shown in FIG. 5B, this implementation of generative decoder 216 may include an input 512 where encoded content frames 120 may be received from another video communication device, another input referred to as profile selection 514 that indicates which of the set of supported codec profiles 116 is to be used for the decoding (described in more detail below), and an output 516 where content frames 118 can be output for rendering and presentation to a user. Generative decoder 216 is also shown to include a series of layers 518 that correspond to the series of layers 508 of generative encoder 214, though the series is reversed. More particularly, a final layer 518-1 corresponding to initial layer 508-1 is shown to be the layer that finally outputs content frames 118 at output 516, and some number of other layers 518-2 through 518-N (corresponding, respectively, to layers 508-2 through 508- N) may also be available for use depending on which codec profile 116 is selected.
[0075] In a mirror image to generative encoder 214 in FIG. 5A, generative decoder 216 of FIG. 5B shows various pipelines corresponding to various codec profiles 116 that may be selected, for decoding a given encoded content frame 120. Opposite to the encoding described above, the decoding may be performed in a pipeline extending from an input of the selected transmission layer (e.g., any of layers 518-N through 518-1 as may be associated with the selected codec profile) to an output of the final layer (i.e., final layer 518-1 in this example). To illustrate the relationship between codec profiles 116 supported by this particular generative decoder 216 implementation and the series of layers 518, FIG. 5B depicts a variety of pipelines, each including different layers 518, for supported codec profiles 116. For example, the first codec profile 116-1 is shown to include a pipeline that is limited to final layer 518-1. In this example, layer 518-1 is therefore not only the final layer but also has been defined as the transmission layer for codec profile 116-1. The second codec profile 116-2 that defines layer 518-2 as the transmission layer is shown to include a pipeline that extends from layer 518-2 to final layer 518-1. The third codec profile 116-3 that defines layer 518-3 as the transmission layer is similarly shown to include a pipeline that extends from layer 518-3 to final layer 518-1. Finally, in the limit, the Nth codec profile 116-N that defines layer 518-N as the transmission layer is shown to include a pipeline that extends fromlayer 518-N to final layer 518-1. In other words, as mentioned above. Nth codec profile 116- N will be understood to include the entire series of layers 518 so as to decode content frames encoded using the entire series of layer 508 (and thereby achieving the lowest bandwidth usage at the cost of the greatest amount of processing).
[0076] Similarly as described above for FIG. 5A, it will be understood that, for a given codec profile 116, layers 518 that come before the selected transmission layer in the series may not be used. For example, while codec profile 116-2 is selected, each encoded content frame 120 may skip layers 518-N through 518-3 to be processed only by layer 518-2 and final layer 518-1 before being output directly as a content frame 118 at output 516. The series of layers 518 may correspond in a one-to-one manner with the series of layers 508 so that generative decoder 216 may accurately unwind whatever extrapolations have been performed by generative encoder 214. Accordingly, it will be understood that layers 518 may be implemented using any of the technologies described above for layers 508 (e.g., CNNs trained using video communication data configured with semantically relevant elements for the video communication application, etc.). Moreover, the number of layers 518 in the series may match the number of layers 508 for a given implementation (e.g., 5 layers, 10 layers, 100 layers, etc ).
[0077] As described above in relation to operations 304, 312, 404, and 412, video communication devices may select codec profiles to use in the encoding and decoding of communication data (e.g., content frames) based on the bandwidth available for communication between the video communication devices and / or based on other conditions. As shown in FIGS. 5A and 5B, this selection may then be input to generative encoder 214 as profile selection 504 and to generative decoder 216 as profile selection 514. The selection of the codec profile in any of these examples may be based at least on the bandwidth available for the video communication devices to exchange data and may be further based on additional factors such as: 1) a processing resource availability of at least one of the first video communication device or the second video communication device; 2) a memoryresource availability of at least one of the first video communication device or the second video communication device; 3) a combination of either or both of these with the bandwidth availability; and / or 4) any other suitable dynamic conditions as may serve a particular implementation.
[0078] To illustrate, FIG. 6 shows illustrative aspects of how a codec profile may be selected by a generative encoder (e.g.. generative encoder 214) in accordance with principles described herein. Specifically, as shown, a process or algorithm labeled as dynamic codecprofile selection 602 is shown to account for at least a bandwidth availability 604, a processing resource availability 606, and a memory resource availability 608. These inputs may be accounted for and weighed in any suitable way. For instance, in certain implementations, bandwidth availability 604 may be accorded the most weight or may even be the only factor considered. When bandwidth availability 604 is relatively low, dynamic codec profile selection 602 may select a codec profile 116 that uses more processing and memory to reduce the amount of data that is transmitted (e.g.. a codec profile 116 that uses a relatively large number of layers). Conversely, when bandwidth availability 604 is relatively high, dynamic codec profile selection 602 may select a codec profile 116 that uses less processing and memory even though that may result in a larger amount of data being transmitted over the network (e.g.. a codec profile 116 that uses a relatively small number of layers).
[0079] While bandwidth availability 604 may be a primary consideration in certain implementations, FIG. 6 shows that it may not be the only consideration that dynamic codec profile selection 602 is configured to account for. For example, in implementations in which compute resources (e.g.. processing and / or memory resources) are relatively sparse and / or network bandwidth happens to be relatively abundant, dynamic codec profile selection 602 may give some or considerable weight also to processing resource availability 606 and / or memory resource availability 608. For example, processing resource availability 606 and memory resource availability 608 may represent resource availability of the video communication device encoding and transmitting content frames, the resource availability of the video communication device receiving and decoding content frames, or a combination of both. In this way, as has been described, video communication devices may efficiently manage the tradeoff between compute resources and network resources that generative codec models described herein may provide.
[0080] Dynamic codec profile selection 602 may be performed by a video communication device (e.g., based on instructions stored in memory 212 and using one or more processors 208) for content frames that are generated by that device and encoded for transmission to the other device. To illustrate. FIG. 6 shows that profile selection 504 may be output from dynamic codec profile selection 602 and that a metadata indication 610 of this profile selection 504 may be included with each encoded content frame 120 that is transmitted (e.g., within a header of the encoded content frames 120, as mentioned above). As has been described, the selection of a codec profile such as profile selection 504 may be indicated or represented in any suitable manner and / or using any suitable data. For instance,the supported codec profiles for a particular implementation may be predetermined and assigned index numbers that profile selection 504 may indicate. As another example, the profile selection 504 for a particular implementation (and the metadata indication 610 indicating the profile selection 504) may indicate an index number assigned to a particular layer of the series of layers (e.g., layers 508) that has been selected as the transmission layer. In either case, profile selection 504 may include data sufficient to indicate to generative encoder 214 which layers 508 are to be used in the encoding pipeline.
[0081] In contrast, FIG. 7 shows illustrative aspects of how a codec profile may be selected by a generative decoder in accordance with principles described herein. Unlike the generative encoder selection shown in FIG. 6, a dynamic codec profile selection 702 performed on the decoder side (e.g.. by generative decoder 216) is shown in FIG. 7 to be based on the metadata indication 610 of profile selection 504. As shown, dynamic codec profile selection 702 inputs this metadata indication 610 that has been made and stored in the metadata of a given encoded content frame 120 and uses that to determine profile selection 514 that sets the profile for generative decoder 216. In this way, it may be ensured that the generative decoder 216 reconstructs each encoded content frame 120 in accordance with the specific layers that were actually used to encode it (regardless of what conditions of the network and / or system may be at the present moment).
[0082] As has been illustrated and described, a generative codec model such as generative codec model 114 may support a plurality of codec profiles 116 that each define a different layer from a series of layers to be the transmission layer. For example, as illustrated and described above in relation to FIGS. 5 A and 5B, a codec profile 116-1 may define a first common layer (i.e., the layer pair including layers 508-1 and 518-1) as the transmission layer, a codec profile 116-2 may define a second common layer (i.e., the layer pair including layers 508-2 and 518-2) as the transmission layer, a codec profile 116-N may define an Nth common layer (i.e., the layer pair including layers 508-N and 518-N) as the transmission layer, and so forth.
[0083] Whichever common layer is defined as the transmission layer, the generative codec model may be trained to ensure that a loss metric corresponding to the codec profile defining that layer satisfies a certain loss threshold. As used herein, a loss metric for a given codec profile may refer to the total loss in detail or quality that may accrue to a content frame when encoded and then decoded using that codec profile. For instance, one loss metric for a codec profile defining an 8thlayer as the transmission layer may represent a maximum discrepancy or loss of detail or quality that could be guaranteed when comparing a rawcontent frame (e.g., a content frame 118-1 generated by video communication device 102-1) to the same content frame after being encoded with an 8-layer pipeline, transmitted across the network, and decoded with a corresponding 8-layer pipeline (e.g., a content frame 118-2 presented by video communication device 102-2). To calibrate and implement a functional generative codec model, each loss metric of a plurality' of loss metrics corresponding to the plurality of codec profiles supported by the generative codec model may be designed and tested to satisfy a particular loss threshold. In this way, the generative codec model may be guaranteed to maintain a particular qualify of sendee regardless of which of the codec profiles are used in the encoding / decoding of any particular content frame.
[0084] To illustrate, FIG. 8 shows illustrative scenarios 802, 804, and 806 in which a generative codec model (e.g., including both a generative encoder 214 and a generative decoder 216) is using different codec profiles that have been selected in accordance with principles described herein. In each of the scenarios illustrated in FIG. 8, an input 502 is shown to include a content frame 118-1 that is to be encoded by generative encoder 214 and transmitted as an encoded content frame 120 on output 506. While transmission through a network (e.g., network 112) will be understood to take place, no network is shown in FIG. 8 (due to space constraints and illustrative simplicity), such that output 506 of generative encoder 214 is also labeled as being input 512 to generative decoder 216 (“506 / 512”). After decoding and reconstructing the encoded content frame 120, generative decoder 216 is then shown to output a content frame 118-2 for presentation to a user as has been described.
[0085] In each different scenario 802, 804, and 806, a profile selection 504 and corresponding profile selection 514 (shown to be connected by dashed lines representing the communication of the profile selection in the metadata, as described above) cause a different codec profile to be used.
[0086] For example, in scenario 802, a first codec profile (e.g., codec profile 116-1) is shown to be selected that designates as the transmission layer a common layer including initial layer 508-1 and final layer 518-1. In other words, in this example, only a single layer may be used both for the encoding and the decoding, thereby minimizing compute resource usage at the expense of network bandwidth usage (e.g.. for a situation where bandwidth may not be particularly scarce).
[0087] In scenario 804, a different codec profile (e.g., codec profile 116-3) is shown to be selected that designates as the transmission layer a common layer including layer 508-3 and layer 518-3. In other words, in this example, three common layers (i.e., layers 508-1 and 518-1, layers 508-2 and 518-2, and layers 508-3 and 518-3) may be used for the encoding andthe decoding, thereby balancing compute resource usage and network bandwidth usage in a different way (e.g.. for a situation where bandwidth may be scarcer than in scenario 802). While scenarios 802 and 806 are show n to represent extremes of how- the tradeoff between compute resources and bandwidth resources may be managed (for minimum compute resource usage in the case of scenario 802 and for minimum bandwidth usage in the case of scenario 806), scenario 804 represents a moderate case where this tradeoff is made away from either extreme. Specifically, in this example, the transmission layer defined by the codec profile is shown to come after an initial layer (e.g., initial layer 508-1) and before a final layer (e.g., layer 508-N) in the series of layers 508. On the decoder side, the transmission layer defined by the selected codec profile will be understood to come after initial layer 518-N and before final layer 518-1.
[0088] In scenario 806, yet another codec profile (e.g., codec profile 116-N) is shown to be selected that designates as the transmission layer a common layer including layer 508-N and layer 518-N. In other words, in this example, N common layers (i.e., all of the layers 508-1 through 508-N and layers 518-N through 518-1) may be used for the encoding and the decoding, thereby minimizing network bandwidth usage at the expense of compute resource usage (e.g., for a situation where bandwidth may be very scarce and / or compute resources available in abundance).
[0089] In each scenario, a respective loss metric is shown as bracketing the common layer associated with the selected transmission layer. Specifically, as shown in scenario 802. a loss metric 812 is depicted between initial layer 508-1 of generative encoder 214 and final layer 518-1 of generative decoder 216. As shown in scenario 804, a loss metric 814 is depicted between layer 508-3 of generative encoder 214 and layer 518-3 of generative decoder 216. As shown in scenario 806, a loss metric 816 is depicted between layer 508-N of generative encoder 214 and layer 518-N of generative decoder 216. As described above, the generative codec model may be trained and otherwise configured (e.g., tested, characterized, etc.) to ensure that each of these loss metrics 812, 814, and 816 (as well as other loss metrics associated with other codec profiles not shown in FIG. 8) satisfies a particular loss threshold. For example, the loss threshold may define the maximum amount of error that is considered acceptable for a transmission of content frames by the system, and the generative codec model may be trained at each layer so that it can be guaranteed that this error is never exceeded regardless of which codec profile may be selected at a given time.
[0090] As has been mentioned, various methods and processes described herein may be implemented at least in part as instructions embodied in a non-transitory computer-readable medium and executable by one or more computing devices. In general, a processor (e.g., a microprocessor) receives instructions, from a non-transitory computer-readable medium (e.g., a memory, etc.), and executes those instructions, thereby performing one or more operations such as the operations described herein. Such instructions may be stored and / or transmitted using any of a variety of known computer-readable media.
[0091] A computer-readable medium (also referred to as a processor-readable medium) includes any non-transitory medium that participates in providing data (e.g., instructions) that may be read by a computer (e.g., by a processor of a computer). Such a medium may take many forms, including, but not limited to, non-volatile media, and / or volatile media. Non-volatile media may include, for example, optical or magnetic disks and other persistent memory. Volatile media may include, for example, dynamic random-access memory (DRAM), which typically constitutes a main memory. Common forms of computer- readable media include, for example, a disk, hard disk, magnetic tape, any other magnetic medium, a compact disc read-only memory (CD-ROM), a digital video disc (DVD), any other optical medium, random access memory (RAM), programmable read-only memory’ (PROM), electrically erasable programmable read-only memory (EPROM), FLASH- EEPROM, any’ other memory chip or cartridge, or any other tangible medium from which a computer can read.
[0092] FIG. 9 shows an illustrative computing system 900 that may be used to implement various devices and / or systems described herein. For example, computing system 900 may include or implement (or partially implement) video communication devices such as video communication devices 102-1 and 102-2, any implementations thereof (e.g., implementation 200 and / or other implementations described herein), any components thereof, and / or other devices used therewith.
[0093] As shown in FIG. 9, computing system 900 may include a communication interface 902, a processor 904, a storage device 906, and an input / output (I / O) module 908 communicatively connected via a communication infrastructure 910. While an illustrative computing system 900 is shown in FIG. 9, the components illustrated in FIG. 9 are not intended to be limiting. Additional or alternative components may be used in other embodiments. Components of computing system 900 shown in FIG. 9 will now be described in additional detail.
[0094] Communication interface 902 may be configured to communicate w ith one or more computing devices. Examples of communication interface 902 include, without limitation, a wired network interface (such as a network interface card), a wireless networkinterface (such as a wireless network interface card), a modem, an audio / video connection, and any other suitable interface.
[0095] Processor 904 generally represents any type or form of processing unit capable of processing data or interpreting, executing, and / or directing execution of one or more of the instructions, processes, and / or operations described herein. Processor 904 may direct execution of operations in accordance with one or more applications 912 or other computerexecutable instructions such as may be stored in storage device 906 or another computer- readable medium.
[0096] Storage device 906 may include one or more data storage media, devices, or configurations and may employ any type, form, and combination of data storage media and / or device. For example, storage device 906 may include, but is not limited to, a hard drive, network drive, flash drive, magnetic disc, optical disc, RAM, dynamic RAM, other non-volatile and / or volatile data storage units, or a combination or sub-combination thereof. Electronic data, including data described herein, may be temporarily and / or permanently- stored in storage device 906. For example, data representative of one or more executable applications 912 configured to direct processor 904 to perform any of the operations described herein may be stored within storage device 906. In some examples, data may be arranged in one or more databases residing within storage device 906.
[0097] I / O module 908 may include one or more I / O modules configured to receive user input and provide user output. One or more I / O modules may be used to receive input for a single virtual experience. I / O module 908 may include any hardware, firmware, software, or combination thereof supportive of input and output capabilities. For example, I / O module 908 may include hardware and / or software for capturing user input, including, but not limited to, a keyboard or keypad, a touchscreen component (e.g., touchscreen display), a receiver (e.g., an RF or infrared receiver), motion sensors, and / or one or more input buttons.
[0098] I / O module 908 may include one or more devices for presenting output to a user, including, but not limited to, a graphics engine, a display (e.g., a display screen), one or more output drivers (e.g., display drivers), one or more audio speakers, and one or more audio drivers. In certain embodiments, I / O module 908 is configured to provide graphical data to a display for presentation to a user. The graphical data may be representative of one or more graphical user interfaces and / or any other graphical content as may sen e a particular implementation.
[0099] The following examples describe implementations of variable-bandwidth video communication using a generative codec model in accordance with principlesdescribed herein.
[0100] Example 1: A method comprising: generating, by a first video communication device at a first scene, a content frame for a visual representation of the first scene; selecting, based on a bandwidth available for the first video communication device to exchange data with a second video communication device at a second scene, a codec profile to be used by a generative codec model, the codec profile defining a transmission layer from a series of layers included in the generative codec model; encoding the content frame, by the generative codec model using the codec profile, in a pipeline extending from an input of an initial layer of the series of layers to an output of the transmission layer; and transmitting, to the second video communication device, the encoded content frame with metadata indicating the transmission layer.
[0101] Example 2: The method of any of the preceding examples, wherein: the content frame is included with an additional content frame generated and transmitted as part of a communication session; the bandwidth on which the selecting of the codec profile is based is a first bandwidth at a first time during the communication session; and the method further compnses: selecting, based on a second bandwidth at a second time during the communication session, an additional codec profile that defines an additional transmission layer from the series of layers, the additional transmission layer different from the transmission layer, encoding the additional content frame, by the generative codec model using the additional codec profile, in an additional pipeline extending from the input of the initial layer to an output of the additional transmission layer, and transmitting, to the second video communication device, the encoded additional content frame with metadata indicating the additional transmission layer.
[0102] Example 3: The method of any of the preceding examples, wherein the selecting the codec profile is further based on a processing resource availability of at least one of the first video communication device or the second video communication device.
[0103] Example 4: The method of any of the preceding examples, wherein the selecting the codec profile is further based on a memory’ resource availability of at least one of the first video communication device or the second video communication device.
[0104] Example 5: The method of any’ of the preceding examples, wherein the generating the content frame for the visual representation of the first scene includes: capturing, from a plurality of vantage points by a plurality' of cameras of the first video communication device, a plurality of images of the first scene; and constructing, based on the plurality of images, a 3D representation of a subject present at the first scene; wherein thevisual representation of the first scene includes the 3D representation of the subject.
[0105] Example 6: The method of any of the preceding examples, wherein the generative codec model is trained using video communication data configured with semantically relevant elements for a video communication application.
[0106] Example 7: The method of any of the preceding examples, wherein: the generative codec model supports a plurality of codec profiles that include the codec profile and that each define a different layer from the series of layers to be the transmission layer; and the generative codec model is trained to ensure that each loss metric of a plurality of loss metrics corresponding to the plurality of codec profiles satisfies a loss threshold.
[0107] Example 8: The method of any of the preceding examples, wherein the generative codec model including the series of layers is implemented as a convolutional neural network including a series of convolutional layers.
[0108] Example 9: The method of any of the preceding examples, wherein the transmission layer defined by the codec profile comes after the initial layer and before a final layer in the series of layers.
[0109] Example 10: A method comprising: receiving, from a first video communication device at a first scene by a second video communication device at a second scene, an encoded content frame for a visual representation of the first scene, the encoded content frame indicating a transmission layer of a series of layers included in a generative codec model; selecting, based on the transmission layer indicated in the encoded content frame, a codec profile to be used by the generative codec model; reconstructing, by the generative codec model using the codec profile, a content frame by decoding the encoded content frame in a pipeline extending from an input of the transmission layer to an output of a final layer of the series of layers; and presenting, by the second video communication device, the content frame.
[0110] Example 11 : The method of any of the preceding examples, wherein: the encoded content frame and an additional encoded content frame are both received as part of a communication session; the additional encoded content frame indicates an additional transmission layer of the series of layers, the additional transmission layer different from the transmission layer; and the method further comprises: selecting, based on the additional transmission layer indicated in the additional encoded content frame, an additional codec profile, reconstructing, by the generative codec model using the additional codec profile, an additional content frame by decoding the additional encoded content frame in an additional pipeline extending from an input of the additional transmission layer to an output of the finallayer of the series of layers, and presenting, by the second video communication device, the additional content frame.
[0111] Example 12: The method of any of the preceding examples, wherein the presenting the content frame includes displaying, on a 3D display of the second video communication device, a 3D representation of a subject present at the first scene.
[0112] Example 13: The method of any of the preceding examples, wherein the generative codec model is trained using video communication data configured with semantically relevant elements for a video communication application.
[0113] Example 14: The method of any of the preceding examples, wherein: the generative codec model supports a plurality of codec profiles that include the codec profile and that each define a different layer from the series of layers to be the transmission layer; and the generative codec model is trained to ensure that each loss metric of a plurality of loss metrics corresponding to the plurality of codec profiles satisfies a loss threshold.
[0114] Example 15: The method of any of the preceding examples, wherein the generative codec model including the series of layers is implemented as a convolutional neural network including a series of convolutional layers.
[0115] Example 16: The method of any of the preceding examples, wherein the transmission layer defined by the codec profile comes after an initial layer in the series of layers and before the final layer.
[0116] Example 17: A video communication device comprising: a memory storing instructions; and one or more processors communicatively coupled to the memory and configured to execute the instructions to perform a process comprising: generating a content frame for a visual representation of a first scene; selecting, based on a bandwidth available for the video communication device to exchange data with an additional video communication device at a second scene, a codec profile to be used by a generative codec model, the codec profile defining a transmission layer from a series of layers included in the generative codec model; encoding the content frame, by the generative codec model using the codec profile, in a pipeline extending from an input of an initial layer of the series of layers to an output of the transmission layer; and transmitting, to the additional video communication device, the encoded content frame with metadata indicating the transmission layer.
[0117] Example 18: The video communication device of any of the preceding examples, further including a machine learning acceleration processor configured to facilitate computation performed by the generative codec model.
[0118] Example 19: A video communication device comprising: a memory storinginstructions; and one or more processors communicatively coupled to the memory and configured to execute the instructions to perform a process comprising: receiving, at a second scene from an additional video communication device at a first scene, an encoded content frame for a visual representation of the first scene, the encoded content frame indicating a transmission layer of a series of layers included in a generative codec model; selecting, based on the transmission layer indicated in the encoded content frame, a codec profile to be used by the generative codec model; reconstructing, by the generative codec model using the codec profile, a content frame by decoding the encoded content frame in a pipeline extending from an input of the transmission layer to an output of a final layer of the series of layers; and presenting, by the additional video communication device, the content frame.
[0119] Example 20: The video communication device of any of the preceding examples, further including a machine learning acceleration processor configured to facilitate computation performed by the generative codec model.
[0120] Various implementations of the systems and techniques described herein can be realized in digital electronic circuitry, integrated circuitry, specially designed ASICs (application specific integrated circuits), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which may be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.
[0121] A number of implementations have been described. Nevertheless, it will be understood that various modifications may be made without departing from the spirit and scope of the description and claims. In addition, the logic flows depicted in the figures do not require the particular order shown, or sequential order, to achieve desirable results. In addition, other steps may be provided, or steps may be eliminated, from the described flows, and other components may be added to, or removed from, the described systems. Accordingly, other implementations are within the scope of the following claims.
[0122] Specific structural and functional details disclosed herein are merely representative for purposes of describing example implementations. Example implementations, however, may be embodied in many alternate forms and should not be construed as limited to only the implementations set forth herein.
[0123] It will be understood that, although the terms first, second, etc. may be used herein to describe various elements, these elements should not be limited by these terms.These terms are only used to distinguish one element from another. A first element could be termed a second element, and, similarly, a second element could be termed a first element, without departing from the scope of the implementations of the disclosure. As used herein, the term and / or includes any and all combinations of one or more of the associated listed items.
[0124] The terminology used herein is for the purpose of describing particular implementations only and is not intended to be limiting of the implementations. As used herein, the singular forms “a,” “an,” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms “comprises,” “comprising,” “includes,” and / or “including,” when used in this specification, specify the presence of the stated features, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, steps, operations, elements, components, and / or groups thereof.
[0125] It will be understood that when an element is referred to as being “coupled,” “connected,” or “responsive” to, or “on,” another element, it can be directly coupled, connected, or responsive to, or on, the other element, or intervening elements may also be present. In contrast, when an element is referred to as being “directly coupled,” “directly connected,” or “directly responsive” to, or “directly on,” another element, there are no intervening elements present. As used herein the term “and / or” includes any and all combinations of one or more of the associated listed items.
[0126] Spatially relative terms, such as “beneath,” “below,” “lower,” “above,” “upper,” and the like, may be used herein for ease of description to describe one element or feature in relationship to another element(s) or feature(s) as illustrated in the figures. It will be understood that the spatially relative terms are intended to encompass different orientations of the device in use or operation in addition to the orientation depicted in the figures. For example, if the device in the figures is turned over, elements described as “below” or “beneath” other elements or features would then be oriented “above” the other elements or features. Thus, the term “below” can encompass both an orientation of above and below. The device may be otherwise oriented (rotated 130 degrees or at other orientations) and the spatially relative descriptors used herein may be interpreted accordingly.
[0127] Unless otherwise defined, the terms (including technical and scientific terms) used herein have the same meaning as commonly understood by one of ordinary skill in the art to which these concepts belong. It will be further understood that terms, such as those defined in commonly used dictionaries, should be interpreted as having a meaning that isconsistent with their meaning in the context of the relevant art and / or the present specification and will not be interpreted in an idealized or overly formal sense unless expressly so defined herein.
[0128] Further to the descriptions above, a user may be provided with controls allowing the user to make an election as to both if and when systems, programs, or features described herein may enable collection of user information (e.g., information about a user's social network, social actions, or activities, profession, a user's preferences, or a user's current location), and if the user is sent content or communications from a server. In addition, certain data may be treated in one or more ways before it is stored or used, so that personally identifiable information is removed. For example, a user's identity' may be treated so that no personally identifiable information can be determined for the user, or a user's geographic location may be generalized, or location information may be obtained (such as to a city', zip code, or state level), so that a particular location of a user cannot be determined. Thus, the user may have control over what information is collected about the user, how that information is used, and what information is provided to the user.
[0129] While certain features of the described implementations have been illustrated as described herein, many modifications, substitutions, changes, and equivalents may occur to those skilled in the art. It is therefore to be understood that the appended claims are intended to cover such modifications and changes as fall within the scope of the implementations. It will be understood that they have been presented by way of example only, not limitation, and various changes in form and details may be made. Any portion of the apparatus and / or methods described herein may be combined in any combination, except mutually exclusive combinations. The implementations described herein can include various combinations and / or sub-combinations of the functions, components, and / or features of the different implementations described. As such, the scope of the present disclosure is not limited to the particular combinations hereafter claimed, but instead extends to encompass any combination of features or example implementations described herein irrespective of whether or not that particular combination has been specifically enumerated in the accompanying claims at this time.
Claims
WHAT IS CLAIMED IS:
1. A method comprising: generating, by a first video communication device at a first scene, a content frame for a visual representation of the first scene; selecting, based on a bandwidth available for the first video communication device to exchange data with a second video communication device at a second scene, a codec profile to be used by a generative codec model, the codec profile defining a transmission layer from a series of layers included in the generative codec model; encoding the content frame, by the generative codec model using the codec profile, in a pipeline extending from an input of an initial layer of the series of layers to an output of the transmission layer; and transmitting, to the second video communication device, the encoded content frame with metadata indicating the transmission layer.
2. The method of claim 1. wherein: the content frame is included with an additional content frame generated and transmitted as part of a communication session; the bandwidth on which the selecting of the codec profile is based is a first bandwidth at a first time during the communication session; and the method further comprises: selecting, based on a second bandwidth at a second time during the communication session, an additional codec profile that defines an additional transmission layer from the series of layers, the additional transmission layer different from the transmission layer, encoding the additional content frame, by the generative codec model using the additional codec profile, in an additional pipeline extending from the input of the initial layer to an output of the additional transmission layer, and transmitting, to the second video communication device, the encoded additional content frame with metadata indicating the additional transmission layer.
3. The method of any of claims 1 to 2, wherein the selecting the codec profile is further based on a processing resource availability of at least one of the first video communication device or the second video communication device.
4. The method of any of claims 1 to 3, wherein the selecting the codec profile is further based on a memory resource availability of at least one of the first video communication device or the second video communication device.
5. The method of any of claims 1 to 4, wherein the generating the content frame for the visual representation of the first scene includes: capturing, from a plurality of vantage points by a plurality of cameras of the first video communication device, a plurality of images of the first scene; and constructing, based on the plurality of images, a 3D representation of a subject present at the first scene; wherein the visual representation of the first scene includes the 3D representation of the subj ect.
6. The method of any of claims 1 to 5, wherein the generative codec model is trained using video communication data configured with semantically relevant elements for a video communication application.
7. The method of any of claims 1 to 6, wherein: the generative codec model supports a plurality of codec profiles that include the codec profile and that each define a different layer from the series of layers to be the transmission layer; and the generative codec model is trained to ensure that each loss metric of a plurality of loss metrics corresponding to the plurality of codec profiles satisfies a loss threshold.
8. The method of any of claims 1 to 7, wherein the generative codec model including the series of layers is implemented as a convolutional neural netw ork including a series of convolutional layers.
9. The method of any of claims 1 to 8, wherein the transmission layer defined by the codec profile comes after the initial layer and before a final layer in the series of layers.
10. A method comprising: receiving, from a first video communication device at a first scene by a second video communication device at a second scene, an encoded content frame for a visual representation of the first scene, the encoded content frame indicating a transmission layer of a series of layers included in a generative codec model; selecting, based on the transmission layer indicated in the encoded content frame, a codec profile to be used by the generative codec model; reconstructing, by the generative codec model using the codec profile, a content frame by decoding the encoded content frame in a pipeline extending from an input of the transmission layer to an output of a final layer of the series of layers; and presenting, by the second video communication device, the content frame.
11. The method of claim 10, wherein: the encoded content frame and an additional encoded content frame are both received as part of a communication session; the additional encoded content frame indicates an additional transmission layer of the series of layers, the additional transmission layer different from the transmission layer; and the method further comprises: selecting, based on the additional transmission layer indicated in the additional encoded content frame, an additional codec profile, reconstructing, by the generative codec model using the additional codec profile, an additional content frame by decoding the additional encoded content frame in an additional pipeline extending from an input of the additional transmission layer to the output of the final layer of the series of layers, and presenting, by the second video communication device, the additional content frame.
12. The method of any of claims 10 to 11, wherein the presenting the content frame includes displaying, on a 3D display of the second video communication device, a 3D representation of a subject present at the first scene.
13. The method of any of claims 10 to 12, wherein the generative codec model is trained using video communication data configured with semantically relevant elements for a video communication application.
14. The method of any of claims 10 to 13, wherein: the generative codec model supports a plurality of codec profiles that include the codec profile and that each define a different layer from the series of layers to be the transmission layer; and the generative codec model is trained to ensure that each loss metric of a plurality of loss metrics corresponding to the plurality of codec profiles satisfies a loss threshold.
15. The method of any of claims 10 to 14, wherein the generative codec model including the series of layers is implemented as a convolutional neural network including a series of convolutional layers.
16. The method of any of claims 10 to 15, wherein the transmission layer defined by the codec profile comes after an initial layer in the series of layers and before the final layer.
17. A video communication device comprising: a memory storing instructions; and one or more processors communicatively coupled to the memory and configured to execute the instructions to perform a process comprising: generating a content frame for a visual representation of a first scene; selecting, based on a bandwidth available for the video communication device to exchange data with an additional video communication device at a second scene, a codec profile to be used by a generative codec model, the codec profile defining a transmission layer from a series of layers included in the generative codec model; encoding the content frame, by the generative codec model using the codec profile, in a pipeline extending from an input of an initial layer of the series of layers to an output of the transmission layer; and transmitting, to the additional video communication device, the encoded content frame with metadata indicating the transmission layer.
18. The video communication device of claim 17, further including a machine learning acceleration processor configured to facilitate computation performed by the generative codec model.
19. A video communication device comprising: a memory storing instructions; and one or more processors communicatively coupled to the memory and configured to execute the instructions to perform a process comprising: receiving, at a second scene from an additional video communication device at a first scene, an encoded content frame for a visual representation of the first scene, the encoded content frame indicating a transmission layer of a series of layers included in a generative codec model; selecting, based on the transmission layer indicated in the encoded content frame, a codec profile to be used by the generative codec model; reconstructing, by the generative codec model using the codec profile, a content frame by decoding the encoded content frame in a pipeline extending from an input of the transmission layer to an output of a final layer of the series of layers; and presenting, by the additional video communication device, the content frame.
20. The video communication device of claim 19, further including a machine learning acceleration processor configured to facilitate computation performed by the generative codec model.
Citation Information
Patent Citations
Image compression and decompression via reconstruction of lower resolution image data
US20220394292A1
Multi-level latent fusion in neural networks for image and video coding
WO2023027873A1
Neural network complexity metric for image processing
WO2023163632A1