Model-based video compression and data transmission
Model-based video compression generates 3D models and transmits residuals to enhance video compression efficiency in bandwidth-limited environments, enabling real-time transmission with improved quality.
Patent Information
- Application Number
- PCT/SG2025/050044
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-24
- Filing Date
- 2025-01-20
- Publication Date
- 2025-07-31
AI Technical Summary
Existing video compression methods are inefficient in bandwidth-limited environments, such as remote areas, disasters, military operations, and underwater scenarios, where traditional compression techniques fail to maintain video quality due to limited communication bandwidth.
Model-based video compression (MBVC) generates a 3D model of the area of interest using techniques like neural radiance fields or 3D Gaussian splatting, synthesizes images based on this model, computes residuals, and transmits latent representations to achieve efficient compression and transmission.
MBVC enables real-time video transmission with improved compression ratios and robustness in bandwidth-limited environments by reducing the bandwidth required for video streaming, particularly in scenarios like underwater operations and remote inspections.
Smart Images

Figure SG2025050044_31072025_PF_FP_ABST
Abstract
Description
MODEL-BASED VIDEO COMPRESSION AND DATA TRANSMISSIONTECHNICAL FIELD
[0001] The present invention relates to methods and systems for model-based video compression (MBVC) and data transmission.BACKGROUND
[0002] Video compression has always been a popular research area, where many traditional and deep video compression methods have been proposed [2], [5], These methods typically rely on signal prediction theory to enhance compression performance by designing high efficient intra and inter prediction strategies and compressing video frames one by one.
[0003] However, there is a need to research on more efficient video compression and transmission while preserving video quality especially in bandwidth limited environments.SUMMARY OF THE INVENTION
[0004] According to a first aspect of the invention, there is provided a method for model-based video compression (MBVC) and data transmission, which includes generating a 3D model of an area of interest based on a first video sequence of the area of interest captured by a camera and a pose of the camera, generating a refined camera pose for a second video sequence of the area of interest, generating a synthesized image of the area of interest based on the generated 3D model of the area of interest and the refined camera pose, comparing the synthesized image with a corresponding image captured by the camera and computing a residual, and transmitting a latent representation of the synthesized image and the residual to a receiving end.
[0005] In some embodiments, the area of interest may be in bandwidth limited environment, and the bandwidth limited environment includes environment with wireless communication channels, space, underwater, Internet of Things (loT), remote areas with limited or no access to high-bandwidth internet, areas with damagedcommunication infrastructure by disasters, and / or areas requiring secure communication channels by military and law enforcement operations.
[0006] In some embodiments, generating the 3D model of the area of interest may include using novel view synthesis (NVS) techniques such as neural radiance fields (NeRF) or 3D Gaussian splatting (3DGS).
[0007] In some embodiments, the method may further includes determining an available communication bandwidth based on communication channel conditions, and dynamically determining an image / video resolution based on the available bandwidth and the generated 3D model of the area of interest.
[0008] In some embodiments, the step of determining the available communication bandwidth may include determining a communication mode between a first communication link with a first bandwidth and a second communication link with a second bandwidth, the first bandwidth being higher than the second bandwidth.
[0009] In some embodiments, the first communication link may include an optical link, and the second communication link may include an acoustic link.
[0010] In some embodiments, the latent representation and the residual may be transmitted via the determined available bandwidth.
[0011] In some embodiments, the method may further include, after the step of generating the 3D model of the area of interest, transmitting the 3D model of the area of interest to the receiving end via a high bandwidth link.
[0012] In some embodiments, the step of generating the 3D model of the area of interest may include capturing the area of interest by the camera to produce the first video sequence of the area of interest, tracking a movement of the camera throughout the sequence and generating accurate pose estimates for each frame, and constructing the 3D model of the area of interest based on captured images and the pose estimates of the camera.
[0013] In some embodiments, dynamically determining the image / video resolution may be based on a machine learning model.
[0014] According to a second aspect of the invention, there is provided a system for model-based video compression (MBVC) and data transmission, which includes a 3D model generator for generating a 3D model of an area of interest based on a first video sequence of the area of interest captured by a camera and a pose of the camera, a motion estimator for generating a refined camera pose for a second video sequence of the area of interest, an image / video Tenderer for generating a synthesized image of the area of interest based on the generated 3D model of the area of interest and the refined camera pose, an encoder for comparing the synthesized image with an image captured by the camera and computing a residual for transmission, and a transmitter for transmitting a latent representation of the synthesized image and the residual to a receiving end.
[0015] In some embodiments, the area of interest may be in bandwidth limited environment, and the bandwidth limited environment includes environment with wireless communication channels, space, underwater, Internet of Things (loT), remote areas with limited or no access to high-bandwidth internet, areas with damaged communication infrastructure by disasters, and / or areas requiring secure communication channels by military and law enforcement operations.
[0016] In some embodiments, the 3D model generator may use novel view synthesis (NVS) techniques such as neural radiance fields (NeRF) or 3D Gaussian splatting (3DGS)
[0017] In some embodiments, the system may further include an adaptive link sense module for determining an available communication bandwidth based on communication channel conditions, and a resolution transformer module for dynamically determining an image / video resolution based on the available bandwidth and the generated 3D model of the area of interest.
[0018] In some embodiments, the adaptive link sense module may be configured to determine a communication mode between a first communication link with a firstbandwidth and a second communication link with a second bandwidth, the first bandwidth being higher than the second bandwidth.
[0019] In some embodiments, the first communication link may include an optical link, and the second communication link may include an acoustic link.
[0020] In some embodiments, the transmitter may transmit the latent representation and the residual via the determined available bandwidth.
[0021] Tn some embodiments, the 3D model generator may be further configured to transmit the generated 3D model of the area of interest to the receiving end via a high bandwidth link.
[0022] In some embodiments, the 3D model generator may be configured to capture the area of interest by the camera to produce the first video sequence of the area of interest, track a movement of the camera throughout the sequence and generate accurate pose estimates for each frame, and construct the 3D model of the area of interest based on captured images and the pose estimates of the camera.
[0023] In some embodiments, the resolution transformer module may include a machine learning model that dynamically adjusts the resolution of a video stream based on the available bandwidth.
[0024] According to a third aspect of the invention, there is provided a method for decoding data compressed by model-based video compression (MBVC), which includes receiving a 3D model of an area of interest based on a first video sequence of the area of interest captured by a camera, receiving a latent representation of a synthesized image of the area of interest based on the generated 3D model of the area of interest and a refined camera pose for a second video sequence of the area of interest, receiving a residual, wherein the residual is computed by comparing the synthesized image with an image captured by the camera, and reconstructing an image / video of the area of interest by incorporating the 3D model of the area of interest and the residual.
[0025] In some embodiments, the generated 3D model of the area of interest may be received via a higher communication link, and the latent representation and the residual may be received via a lower communication link.
[0026] According to a fourth aspect of the invention, there is provided a system for decoding data compressed by model-based video compression (MBVC), which includes a model based decoder. The model based decoder is configured to receive a 3D model of an area of interest based on a first video sequence of the area of interest captured by a camera, receive a latent representation of a synthesized image of the area of interest based on the generated 3D model of the area of interest and a refined camera pose for a second video sequence of the area of interest, receive a residual, wherein the residual is computed by comparing the synthesized image with an image captured by the camera, and reconstruct an image / video of the area of interest by incorporating the 3D model of the area of interest and the residual.
[0027] In some embodiments, the model based decoder may be configured to receive the generated 3D model of the area of interest via a higher communication link, and to receive the latent representation and the residual via a lower communication link.
[0028] According to a fifth aspect of the invention, there is provided a system for inspection of underwater environment, which includes the system for MBVC and data transmission of the second aspect, an unmanned underwater vehicle (UUV) in communication with the system for MBVC and data transmission, and a base station for establishing a wireless communication link with the UUV and a tethered link to communicate with a remote station.
[0029] In some embodiments, the UUV may include a hybrid underwater vehicle (HUV).
[0030] In some embodiments, the HUV may establish a communication link for virtual tethering via an acoustic link for high frequency information and an optical link for short-range low latency communication.
[0031] In some embodiments, the UUV may be equipped with the camera.
[0032] According to a sixth aspect of the invention, there is provided a system for model-based video compression (MB VC) and data transmission, which includes one or more processors, and memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for performing or facilitating performing of the method of the first aspect.
[0033] According to a seventh aspect of the invention, there is provided a system for decoding data compressed by model-based video compression (MBVC), which includes one or more processors, and memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for performing or facilitating performing of the method of the third aspect.
[0034] According to an eighth aspect of the invention, there is provided a non- transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors, the one or more programs including instructions for performing or facilitating performing of the method of the first aspect or the third aspect.
[0035] Other features and aspects of the invention will become apparent by consideration of the detailed description and accompanying drawings. Any feature(s) described herein in relation to one aspect or embodiment may be combined with any other feature(s) described herein in relation to any other aspect or embodiment as appropriate and applicable.BRIEF DESCRIPTION OF DRAWINGS
[0036] Embodiments of the invention will now be described, by way of example, with reference to the accompanying drawings in which:
[0037] Fig. 1 shows a block diagram illustrating a framework of model-based video compression and data transmission according to an embodiment of the invention.
[0038] Fig. 2 shows a schematic diagram of an architecture for virtual tethering of a remotely operated vehicle according to an embodiment of the invention.
[0039] Fig. 3 shows an example system for inspection of underwater environment according to an embodiment of the invention.
[0040] Before any embodiments of the invention are explained in detail, it is to be understood that the invention is not limited in its application to the details of embodiment and the arrangement of components set forth in the following description or illustrated in the following drawings. The invention is capable of other embodiments and of being practiced or of being carried out in various ways. Also, it is to be understood that the phraseology and terminology used herein is for the purpose of description and should not be regarded as limiting.DETAILED DESCRIPTION
[0041] Hereinafter, some embodiments of the invention will be described in detail with reference to the drawings.
[0042] Efficient Model-Based Video Compression (MBVC) and Data Transmission in Bandwidth Limited Environments
[0043] Model-based video compression (MBVC) is a novel video compression technique that leverages a 3D model of the video scene to reduce redundancy in the video data. Unlike traditional video compression techniques, which rely on transform coding and motion compensation, MBVC works by first creating a 3D model of the operating environment. This model is then used to synthesize a frame of the scene, given the location of the camera in the operating environment The difference between the synthesized frame and the actual camera frame is then compressed and transmitted, resulting in significant bandwidth savings. MBVC has the potential to enable real-time video transmission even in the most bandwidth limited environments.
[0044] The bandwidth required for streaming HD, FHD, and UHD video content depends on various factors such as resolution, codec, bitrate, frame rate, and compression efficiency. Using the codec, for example, H.265 / HEVC, the minimum bandwidth requirements vary anywhere between 1Mbps and 20Mbps, depending on video resolution and frame rate
[0001] , [2]. There are numerous scenarios wherecommunication bandwidth is a precious resource and needs to be used efficiently. Here are some examples:
[0045] - Remote areas: Rural areas and mountainous regions may have limited or no access to high-bandwidth internet
[0046] - Disasters: Disasters, such as earthquakes, hurricanes, and floods, can damage or destroy communication infrastructure, resulting in limited bandwidth.
[0047] - Military and law enforcement: Military and law enforcement operations may require secure communication channels with limited bandwidth.
[0048] - Space exploration: Spacecraft operating in deep space must communicate with Earth over long distances, which can limit bandwidth.
[0049] - Internet of Things (loT): loT devices, such as sensors and actuators, often have limited bandwidth due to their small size and power constraints.
[0050] In addition to the above scenarios, communication bandwidth may also be limited by the following factors:
[0051] - The type of communication channel: Wireless communication channels, such as cellular networks and satellite links, typically have lower bandwidth than wired communication channels, such as fiber optic cables.
[0052] - The medium of wireless communication channel: Wireless communication underwater is significantly more challenging than land-based communication due to the severe attenuation of electromagnetic signals. Even in clear water, light can only penetrate a few tens of meters, and much less in water that is turbid due to suspended sediments. While acoustic systems can communicate over long ranges, they have limited data rates and significant latency due to the speed of sound in water. Table 1 below summarizes some example underwater communication modalities in practice in the context of unmanned underwater vehicle (UUV) operations.TABLE 1 Communication Method Data Rate Operating RangeTethered (Conventional / Fiber optic cables) few tens of Gbps 0 - 20kmCommercial off-the-shelf Acoustic modem [3] 1 - 65kbps 0 - 15kmCommercial off-the-shelf Optical Modem |4| I - 3Mbps
[0042] Fig. 1 shows a block diagram illustrating a framework of MBVC and data transmission according to an embodiment of the invention. MBVC is a video compression technique that leverages the ability to build a model of the operating environment in the first pass to achieve significant improvement in compression ratios and robustness in the subsequent passes. The block diagram in Fig. 1 illustrates the fundamental components of model-based video compression.
[0043] In the First Pass, a camera captures a video sequence of an operating environment (i.e., an area of interest). Then, an ego-motion estimator tracks the camera's movement throughout the sequence, generating accurate pose estimates for each frame. A 3D model generator uses the camera poses and captured images to construct a detailed 3D model of the environment
[0044] In the Subsequent Passes, for each frame in the video sequence, a fine-scale motion estimator generates a refined camera pose estimate using the current image and the camera's coarse location. An adaptive link sense module senses the operating environment and decides a communication mode. This in-turn determines an available communication bandwidth. A resolution transformer module dynamically determines an image / video resolution based on the available bandwidth and the 3D model / digital twin of the operating environment. Equipped with 3D model / digital twin from the first pass and refined pose, an image / video Tenderer generates a synthesized image of the operating environment that the camera is expected to capture. The information required to generate the synthesized image can be referred to as latent representation, that may encompass pose and lighting information, amongst others. A model-based encoder then compares the synthesized image with the actual camera-captured image and computes a residual, which is typically very small due to the low entropy of the residual. The residual, along with the latent representation, is then transmitted wirelessly to a receiver.
[0045] At the Receiving End, a model-based decoder leverages the received residual and the 3D model to reconstruct the original camera-captured image. A resolution transformer can be effectively a super-resolver.
[0046] Meanwhile, repetitive offshore underwater inspection operations using UUVs are often conducted in remote and bandwidth-constrained environments. Recent advances in acoustic and optical wireless underwater communications, coupled with efficient MBVC systems, have the potential to significantly improve offshore underwater inspection operations using unmanned underwater vehicles by enabling realtime wireless video transmission in low-bandwidth environments. The next sub-section will describe remote human-m-the-loop offshore inspection operations using UUVs, which incorporates a use of MBVC.
[0047] Virtual Tethering in Bandwidth Limited Environments: Remote human- m-the-loop offshore inspection operations using UUVs
[0048] In the subsea domain, UUVs are categorized into three major classes: remotely operated, autonomous and hybrid. Remotely operated vehicles (ROVs) have become an indispensable technology for offshore operations, primarily for asset integrity management. The tether provides power and communication links between the ROV and its operator. However, the tether is also the key drawback of this technology, necessitating a large surface support vessel, complex tether management systems limiting operations in high sea states, increasing operational risk due to potential entanglement, and preventing the use of multiple ROVs in the same operational area. Currently, most industrial ROV operations are manually controlled, with limited or no automatic control functions or autonomy. Thus, operational efficiency is highly dependent on the experience of the human operator. Autonomy in ROV operations is a stepping stone towards increasing efficiency, reducing operational expenditure (OPEX), and improving overall health, safety, and environment (HSE) standards.
[0049] Autonomous underwater vehicles (AUVs), which rely on onboard power supplies and intelligent control and navigation units, have become recognized tools for seafloor mapping, surveys, and mine countermeasure operations. However, acoustic communication is slow, has high latency, and is notoriously unreliable. Without real-time feedback to the human operator, AUVs have been mostly restricted to mapping and survey operations, with almost no use in inspection and intervention operations. When operating close to structures or when manipulation is required, the risk is too high for an operator to trust an AUV to make critical decisions autonomously, and the low communication bandwidth and reliability do not provide the means for close supervision or human-robot collaboration in decision-making.
[0050] Hybrid underwater vehicles (HUVs) are an emerging class of UUVs that combines the best features of remotely RO Vs and AUVs. HUVs are untethered, but have high-speed communication capabilities over short ranges.
[0051] With the development of key technology components in the areas of 3D computer vision-based digital twinning and model-based video compression, coupled with significant advances in robotics, automation, sensors, and communication, it is possible to perform a range of undersea tasks using untethered underwater vehicles under careful human supervision. It is possible to radically redefine the meaning of the words “tethered system” through a hybrid approach by including virtual tethering via high frequency acoustic and / or optical communication at a short range, and with an extremely small optical fiber base station providing remote surface connectivity. By collaborating with an onshore / offshore remote human operator and a digital twin of the environment, coupled with MBVC systems, tetherless hybrid underwater vehicles (HUVs) will be able to perform a range of tasks at a remote location that might otherwise require a conventional ROV and a support ship on site.
[0052] An architecture for virtual tethering of ROVs according to some embodiments of the invention is described in detail hereinafter.
[0053] A Remotely Operated Vehicle (ROV) plays a crucial role in facilitating underwater inspections and interventions in offshore operations. The use of tether in ROVs usually requires a significant on-site infrastructure such as a large tether management system (TMS), a support vessel with ROV operators aboard However, the tether is indispensable for enabling operators to maintain comprehensive control over the vehicle’s operation but remains susceptible to entanglement with subsea obstacles.Furthermore, the thrust required to carry wired tethers increases rapidly in proportion to length of the tether.
[0054] Autonomous Underwater Vehicles can operate without a tether and thus be used in some off-shore operations such as a site survey or scanning the seabed. However, human-in-the-loop control is still required for executing intricate subsea tasks. Manipulation to control valves or precise maneuvering in proximity to subsea infrastructure are some examples of these tasks. This requirement is the strong motivation behind developing Hybrid ROVs (HROVs), a remotely operated vehicle without a wired connection but still retaining the operational benefits inherent to tethered systems.
[0055] An early attempt in the development of HROVs involves replacing the wired tether with acoustic links
[0018] , The main challenge in using acoustic modems as compared to a wired connection is the lower data throughput. The data rates of an acoustic modem are not sufficient for real-time video streaming and achieve fewer Frames Per Second (FPS) when using off-the-shelf video compression algorithms. This approach may indeed offer wireless connectivity but falls short in delivering real-time video to operators due to inherent limitations in acoustic communication.
[0056] Optical modems can be integrated with acoustic modems to overcome the throughput limitations
[0019] , This integration provides a low-latency wireless connection between a base station and an underwater vehicle. The combination of these technologies is also used to demonstrate how ROVs can be operated without requiring a wired tether
[0020] , In this demonstration, the optical modem facilitates real-time video transmission and the acoustic modem is used for controlling the manipulator. The combination of modems offers a high throughput wireless link but within a limited operational radius, as the throughput of optical modems diminishes with increasing turbidity and thus substantially limits the operational area of ROVs.
[0057] Replacing the wired tether with an alternative that provides a large operational area and responsive control is still an open problem. Using the recent advancements in machine learning, some embodiments of the invention propose anarchitecture that works as a virtual tether and has the operational benefits similar to having a wired connection, thus facilitating remote human-in-the-loop wireless operations.
[0058] Method according to Some Embodiments
[0059] The offshore inspection and intervention missions using ROVs are typically executed over the same geographical area in successive runs. In the first inspection run, the ROV is equipped with additional instrumentation such as acoustic modems for positioning and tethers for power, control commands and communication. With these equipment, an ROV can collect as much information as possible to characterize the baseline state of the survey area and to inspect the target infrastructure as thoroughly as possible. The collected data normally includes visual and sonar profiles and navigational information.
[0060] It is possible to use this collected data to train machine learning models that help minimize data that has to be sent over the tether during successive runs. For instance, the utilization of autoencoder-based compression for scientific data has demonstrated compression ratios of up to 50%
[0021] . However, even with this degree of compression, it remains infeasible for acoustic modems to provide a high FPS video stream. As a result, this simple architecture using an environment specific compression and decompression model can provide a virtual tether with large operational area but with lower FPS In order to achieve even higher FPS, some embodiments of the invention employ a novel machine learning based video compression architecture that capitalizes on repetitive nature of offshore missions to transmit only the novel information in real time. This enables to achieve good video quality and FPS even with relatively low bandwidth afforded by the acoustic communication link. The machine learning based models used for video compression are compute intensive, but it is demonstrated that modern-day embedded computing capability has reached a point where this technique has become practical.
[0061] According to some embodiments of the invention, an optical modem is integrated alongside an acoustic modem for opportunistic low-latency short range communication. It can be automatically switched between acoustic and opticalcommunication, leveraging higher rate optical communication link when operating conditions permit it. The optical link can also be used for downloading large files such as mission logs. This complete architecture of transmitting only novel information and the combination of optical and acoustic modem provides a viable virtual tether for ROV operations. Fig. 2 shows a graphical depiction of this architecture.
[0062] Experiments and Results
[0063] The implementation of the architecture may require retrofitting an optical and an acoustic modem onto a commercially available ROV, and install an additional Single Board Computer (SBC) to run the virtual tether architecture. Experiments are conducted in an artificial ocean basin
[0022] , where the ROV is used to execute underwater structure inspections. The first inspection mission is performed to collect data and implement the virtual tether architecture. Subsequent missions use the virtual tether architecture for data compression, and it is observed that throughput requirements are significantly reduced. For example, a 720x320 pixel images can be transmitted using only 16 to 35 kilobits after applying the machine-learning based compression techniques according to the embodiments of the invention. 1 to 2 FPS video transmission is demonstrated with during the experiment. This can serve as a proof-of-concept for the proposed system, and a baseline for further optimization to achieve higher FPS video transmission in the future.
[0064] System for Inspection of Underwater Environment
[0065] As previously described, MBVC is a novel approach to video compression that leverages the ability to build a model of the operating environment in the first pass to achieve significant compression and robustness in the subsequent passes. Unlike traditional techniques, which exploit spatial and temporal redundancy in video content, MBVC harnesses the power of machine learning to predict model frames given a prebuilt model and pose. This enables highly efficient encoding and decoding, while preserving video quality even in challenging environments.
[0066] The repetitive nature of offshore inspection missions renders remote human- in-the-loop offshore inspection operations using hybrid underwater vehicles (HUVs) acompelling use case for MBVC. MBVC leverages the repetitive nature of these offshore inspection missions by transmitting only the residual, the difference between the model- synthesized image and the camera-captured image. This can significantly reduce the bandwidth required to transmit video from HUVs to remote human operators, enabling more efficient and effective inspection operations. To understand how this is achieved, the concept of operations is outlined below with reference to Fig. 3.
[0067] Step 1 : An HUV equipped with a camera and navigation pay load is deployed in the vicinity of a structure to be inspected.
[0068] Step 2: The HUV autonomously maps the structure that is being inspected, using photogrammetry, sensor-fusion, mapping, and localization technologies.Photorealistic multi-sensor 3D models of the structure are then reconstructed offline. In an embodiment, the 3D model / digital twin of the operating environment can be built using novel view synthesis (NVS) techniques such as neural radiance fields (NeRF) or 3D Gaussian splatting (3DGS).
[0069] Step 3 : The 3D model is inspected to determine if a closer look or intervention is required. Assisted by automated data analysis and machine learning technologies, an inspector or an operator will be able to quickly and reliably determine if:
[0070] - an anomaly is detected during inspection and needs a closer look with a different sensor,
[0071] - a closer look into certain parts of the structure is needed, or
[0072] - certain light interventions need to be performed on the structure.
[0073] Step 4: If closer inspection or light intervention is required, the HUV is deployed, again. It navigates to the area of interest autonomously and, if required, transitions to remote-operated mode, allowing an operator to control or supervise further operations. The HUV combines the tetherless mobility of an autonomous underwater vehicle with the remote operability of a remotely operated vehicle, enabling it to cover a large area without the risk of entanglement. To further reduce reliance on surface support, an ad-hoc underwater base station is deployed alongside the HUV. Oncedeployed, the base station establishes a high-speed short-range underwater wireless communication link with the HU V and a tethered link to a surface buoy, which relays communications to a remote onshore ground station. An adaptive link sense module determines whether to use a high-bandwidth optical link or a low-bandwidth acoustic link, depending on the channel conditions. Equipped with digital twin of the operating environment, a resolution transformer decides on the image quality based on the available bandwidth.
[0074] Step 5: Leveraging a 3D model / digital twin generated in step 2, the image / video Tenderer generates a synthesized image based on the latent representation of the scene that the onboard HUV camera is expected to capture. The model-based encoder then compares the synthesized image with the onboard HUV camera-captured image and computes the residual, which is transmitted. Given the repetitive nature of these offshore inspection and intervention missions, the residual images / videos are significantly smaller in size than the onboard HUV images / videos, thus significantly reducing the bandwidth required to transmit image / video from the HUV to a remote station (or remote human operators).
[0075] Step 6: At the onshore operator end, a model-based decoder leverages the received residual image and the 3D model generated in step 2 to reconstruct the original onboard HUV camera-captured image. The resolution transformer at the receiving end effectively acts as a superresolver.
[0076] Conclusion
[0077] In conclusion, model-based video compression is an emerging video compression technique that leverages the ability to build 3D models of the operating environment to achieve significant improvements in compression ratios and robustness, especially in scenarios where the site is revisited regularly. This renders MBVC an ideal solution for a diverse range of applications, including remote offshore inspection operations using unmanned underwater vehicles.
[0078] Example features of some embodiments
[0079] Model Based Video Compression
[0080] Video compression has always been a popular research area, where many traditional and deep video compression methods have been proposed [2], [5], These methods typically rely on signal prediction theory to enhance compression performance by designing high efficient intra and inter prediction strategies and compressing video frames one by one. Model-based video compression (MBVC) is a novel approach to video compression that leverages the ability to build a model of the operating environment in the first pass to achieve significant improvement in compression ratios and robustness in the subsequent passes. Unlike traditional techniques, which exploit spatial and temporal redundancy in video content, MBVC harnesses the power of machine learning to predict model frame given a pre-built 3D model / digital twin. This enables highly efficient encoding and decoding, while preserving video quality even in challenging environments.
[0081] 3D Model Generation using Novel View Synthesis (NVS)
[0082] Neural radiance fields (NeRFs) and 3D Gaussian splatting (3DGS) are an emerging set of techniques employing machine learning techniques to 3D scene representation and novel view synthesis (NVS) [6], NeRFs represent a 3D scene as a continuous function that predicts the color and density of light at any point in space. 3DGS represent a 3D scene as a combination of Gaussian shapes of various colors. These representations are learned from a set of 2D images of the scene, taken from multiple viewpoints. They have wide range of potential applications ranging from augmented reality, autonomous driving, 3D printing to medical imaging
[0083] However most NeRF and 3DGS implementations are land-based and don't account for backs cattering from the medium (e.g., underwater) or lighting variations (e.g., ambient light or artificial illumination) with an exception, where color restoration of an underwater scenes is attempted [7], Some embodiments of the invention propose the use of eRF or 3DGS based 3D models generated from monocular cameras mounted on unmanned underwater vehicles as priors for model-based video compression
[0084] Adaptive Link Sense Module
[0085] The adaptive link sense module is a software module that is developed as part of the Unet underwater communication framework
[0011] , Its unique ability to sense the communication channel and determine the appropriate communication mode and video encoding parameters maximizes the quality of service (QoS) in the context of model-based video compression (MBVC).
[0086] The adaptive link sense module can provide the ability to switch between communication modes based on the channel conditions and adaptively adjust video encoding parameters, in the context of MBVC, to maximize the QoS.
[0087] Resolution Transformer
[0088] The resolution transformer is a machine learning model that dynamically adjusts the resolution of a video stream based on the available bandwidth, making it ideal for reducing the bandwidth required to transmit video over limited bandwidth channels, such as underwater wireless communication links. The adaptive link sense module provides the resolution transformer with the available channel bandwidth, which can vary depending on the operating environment.
[0089] At the transmitting end, the resolution transformer analyzes the video stream and the pre-existing 3D model to generate a lower-resolution version of the video stream. The model-based encoder then generates a residual signal that represents the difference between the lower-resolution video stream and the synthesized video stream. The residual signal along with the latent representation is then transmitted, which significantly reduces the bandwidth required compared to transmitting the original video stream. At the receiving end, the model-based decoder decodes the received residual signal, and the resolution transformer aided by pre-existing 3D model reconstructs the original video stream, effectively acting as a superresolver.
[0090] Example applications of some embodiments
[0091] The MBVC has the potential to enable real-time wireless video transmission in low-bandwidth underwater environments. This may allow for remote onshoreoperations to manage offshore assets using UUVs. This would particularly reduce OPEX in the following sectors:
[0092] - Offshore asset integrity management (e g., Oil rigs, mooring lines, structures)
[0093] - Offshore renewables (e g., Offshore Wind-farm management)
[0094] - Aquaculture
[0095] - Port surveillance
[0096] Example features and advantages of some embodiments
[0097] The MBVC according to some embodiments of the invention can provide real-time video transmission in low bandwidth scenarios, higher quality of service (QoS), and improved robustness to errors.
[0098] According to some embodiments of the invention, there is provided the NeRF or 3DGS based 3D models generated from monocular cameras mounted on unmanned underwater vehicles as priors for MBVC. This can allow photorealistic novel view synthesis of underwater structures in turbid waters with variable illumination. The synthesized image can be varied in appearance (exposure, lighting, etc.) without affecting 3D geometry.
[0099] According to some embodiments of the invention, there is provided the adaptive link sense module which allows adaptive switching between communication modes and adjusting video encoding parameters based on channel conditions. This can maximize channel dependent quality of service (QoS).
[0100] According to some embodiments of the invention, the resolution transformer allows dynamic video / image stream resolution adjustment based on environment specific channel conditions by using pre-existing 3D models. This can provide improved quality of video using information from the operating environment, and improved compression
[0101] It will be appreciated by a person skilled in the art that variations and / or modifications may be made to the described and / or illustrated embodiments of the invention to provide other embodiments of the invention. The described / or illustrated embodiments of the invention should therefore be considered in all respects as illustrative, not restrictive. Example optional features of some embodiments of the invention are provided in the summary and the description. Some embodiments of the invention may include one or more of these optional features. Some embodiments of the invention may lack one or more of these optional features.
[0102] REFERENCES
[0103] All referenced literatures throughout this disclosure are incorporated herein by reference in their entirety, which include the following references:[1] “Broadband Speed Guide.” Accessed: Oct. 26, 2023. [Online], Available: https: / / www.fcc.gov / consumers / guides / broadband-speed-guide[2] J. Klink, “A Method of Codec Comparison and Selection for Good Quality Video Transmission Over Limited-Bandwidth Networks,” Sensors (Basel), vol. 21 , no. 13, p. 4589, Jul. 2021 , doi: 10.3390 / s21134589.[3] M. Y. I. Zia, J. Poncela, and P. Otero, “State-of-the-Art Underwater Acoustic Communication Modems: Classifications, Analyses and Design Challenges,” Wireless Pers Commun, vol. 116, no. 2, pp. 1325-1360, Jan. 2021 , doi: 10.1007 / s11277-020- 07431-x.[4] P. Leon et al., “A new underwater optical modem based on highly sensitive Silicon Photomultipliers,” in OCEANS 2017 - Aberdeen, Aberdeen, United Kingdom: IEEE, Jun. 2017, pp. 1-6. doi: 10.1 109 / OCEANSE.2017.8084586.[5] T. M. Hoang and J. Zhou, “Recent trending on learning based video compression: A survey,” Cognitive Robotics, vol. 1 , pp. 145-158, Jan. 2021 , doi: 10.1016 / j.cogr.2021.08.003.[6] A. Tewari et al., “Advances in Neural Rendering.” arXiv, Mar. 30, 2022. Accessed: Oct. 1 1 , 2022. [Online]. Available: http: / / arxiv.org / abs / 2111.05849[7] D. Levy et al., “SeaThru-NeRF: Neural Radiance Fields in Scattering Media,” in 2023 IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vancouver, BC, Canada: IEEE, Jun. 2023, pp. 56-65. doi:10.1109 / CVPR52729.2023.00014.[8] M. Tancik et al., “Block-NeRF: Scalable Large Scene Neural View Synthesis.” arXiv, Feb. 10, 2022. doi: 10.48550 / arXiv.2202.05263.[9] W. F. Low and G. H. Lee, “Robust e-NeRF: NeRF from Sparse & Noisy Events under Non-Uniform Motion.” arXiv, Sep. 15, 2023. Accessed: Nov. 01 , 2023. [Online], Available: http: / / arxiv.org / abs / 2309.08596
[0010] R. Martin-Brualla, N. Radwan, M. S. M. Sajjadi, J. T. Barron, A. Dosovitskiy, and D. Duckworth, “NeRF in the Wild: Neural Radiance Fields for Unconstrained Photo Collections,” in 2021 IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR), Nashville, TN, USA: IEEE, Jun. 2021 , pp. 7206-7215. doi:10.1109 / CVPR46437.2021 .00713.
[0011] “UnetStack.” Accessed: Nov. 01 , 2023. [Online], Available: https: / / unetstack.net
[0012] M. Doniec, I. Topor, M. Chitre, and D. Rus, “Autonomous, Localization-Free Underwater Data Muling Using Acoustic and Optical Communication,” in Experimental Robotics: The 13th International Symposium on Experimental Robotics, J. P. Desai, G. Dudek, O. Khatib, and V. Kumar, Eds., in Springer Tracts in Advanced Robotics. , Heidelberg: Springer International Publishing, 2013, pp. 841-857. doi: 10.1007 / 978-3- 319-00065-7_56.
[0013] N. Farr, A. Bowen, J. Ware, C. Pontbriand, and M. Tivey, “An integrated, underwater optical / acoustic communications system,” in OCEANS’10 IEEE SYDNEY, Sydney, Australia: IEEE, May 2010, pp. 1-6. doi: 10.1 109 / OCEANSSYD.2010.5603510.
[0014] E. P. Aswathy K. Cherian, “A Novel AlphaSRGAN for Underwater Image Super Resolution,” Computers, Materials & Continua, vol. 69, no. 2, pp. 1537-1552, 2021 , doi: 10.32604 / cmc.2021.018213.
[0015] M. J. Islam, S. Sakib Enan, P. Luo, and J. Sattar, “Underwater Image SuperResolution using Deep Residual Multipliers,” in 2020 IEEE International Conference on Robotics and Automation (ICRA), 2020, pp. 900-906. doi:10.1109 / ICRA40945.2020.9197213.
[0016] H. Wang et al., “Simultaneous restoration and super-resolution GAN for underwater image enhancement,” Frontiers in Marine Science, vol. 10, 2023, doi: 10.3389 / fmars.2023.1162295.
[0017] Y. Lai et al. , “Super resolution of underwater image based on generating generative adversarial networks,” in AOPC 2022: Atmospheric and Environmental Optics, J. Liu, L. Xu, and J. Gao, Eds., SPIE, 2023, p. 125610E. doi: 10.1117 / 12.2652062.
[0018] R. Dunbar and A. Settery, “Video communications for the untethered submersible rover,” in Proceedings of the 1985 4th International Symposium on Unmanned Untethered Submersible Technology, vol. 4. IEEE, 1985, pp. 140-149.
[0019] N. Farr, A. Bowen, J. Ware, C. Pontbriand, and M. Tivey, “An integrated, underwater optical / acoustic communications system,” in OCEANS’ 10 IEEE SYDNEY. IEEE, 2010, pp. 1-6.
[0020] A. D. Bowen, M. V. Jakuba, N. E. Farr, J. Ware, C. Taylor, D. Gomez-Ibanez, C. R. Machado, and C. Pontbriand, “An un-tethered rov for routine access and intervention in the deep sea,” in 2013 oceans-san diego. IEEE, 2013, pp. 1-7.
[0021] T. Liu, J. Wang, Q. Liu, S. Alibhai, T. Lu, and X. He, “High-ratio lossy compression: Exploring the autoencoder to compress scientific data,” IEEE Transactions on Big Data, 2021.
[0022] “TCOMS Research & Development.” [Online]. Available: https: / / www.tcoms.sg / research-development /
Claims
CLAIMS1. A method for model-based video compression (MBVC) and data transmission, comprising: generating a 3D model of an area of interest based on a first video sequence of the area of interest captured by a camera and a pose of the camera; generating a refined camera pose for a second video sequence of the area of interest; generating a synthesized image of the area of interest based on the generated 3D model of the area of interest and the refined camera pose; comparing the synthesized image with a corresponding image captured by the camera and computing a residual; and transmitting a latent representation of the synthesized image and the residual to a receiving end.
2. The method of claim 1, wherein the area of interest is in bandwidth limited environment, and the bandwidth limited environment comprises environment with wireless communication channels, space, underwater, Internet of Things (ToT), remote areas with limited or no access to high-bandwidth internet, areas with damaged communication infrastructure by disasters, and / or areas requiring secure communication channels by military and law enforcement operations.
3. The method of claim 1, wherein generating the 3D model of the area of interest comprises using novel view synthesis (NVS) techniques such as neural radiance fields (NeRF) or 3D Gaussian splatting (3DGS)4. The method of claim 1, further comprising: determining an available communication bandwidth based on communication channel conditions; and dynamically determining an image / video resolution based on the available bandwidth and the generated 3D model of the area of interest.
5. The method of claim 4, wherein the step of determining the available communication bandwidth comprises determining a communication mode between a first communication link with a first bandwidth and a second communication link with a second bandwidth, the first bandwidth being higher than the second bandwidth.
6. The method of claim 5, wherein the first communication link comprises an optical link, and the second communication link comprises an acoustic link.
7. The method of claim 4, wherein the latent representation and the residual is transmitted via the determined available bandwidth.
8. The method of claim 1, further comprising, after the step of generating the 3D model of the area of interest, transmitting the 3D model of the area of interest to the receiving end via a high bandwidth link.
9. The method of claim 1, wherein the step of generating the 3D model of the area of interest comprises: capturing the area of interest by the camera to produce the first video sequence of the area of interest; tracking a movement of the camera throughout the sequence and generating accurate pose estimates for each frame; and constructing the 3D model of the area of interest based on captured images and the pose estimates of the camera.
10. The method of claim 4, wherein dynamically determining the image / video resolution is based on a machine learning model.
11. A system for model-based video compression (MBVC) and data transmission, comprising:a 3D model generator for generating a 3D model of an area of interest based on a first video sequence of the area of interest captured by a camera and a pose of the camera; a motion estimator for generating a refined camera pose for a second video sequence of the area of interest; an image / video Tenderer for generating a synthesized image of the area of interest based on the generated 3D model of the area of interest and the refined camera pose; an encoder for comparing the synthesized image with an image captured by the camera and computing a residual for transmission; and a transmitter for transmitting a latent representation of the synthesized image and the residual to a receiving end.
12. The system of claim 11, wherein the area of interest is in bandwidth limited environment, and the bandwidth limited environment comprises environment with wireless communication channels, space, underwater, Internet of Things (loT), remote areas with limited or no access to high-bandwidth internet, areas with damaged communication infrastructure by disasters, and / or areas requiring secure communication channels by military and law enforcement operations.
13. The system of claim 11, wherein the 3D model generator uses novel view synthesis (NVS) techniques such as neural radiance fields (NeRF) or 3D Gaussian splatting (3DGS).
14. The system of claim 11, further comprising: an adaptive link sense module for determining an available communication bandwidth based on communication channel conditions; and a resolution transformer module for dynamically determining an image / video resolution based on the available bandwidth and the generated 3D model of the area of interest;15. The system of claim 14, wherein the adaptive link sense module is configured to determine a communication mode between a first communication link with a firstbandwidth and a second communication link with a second bandwidth, the first bandwidth being higher than the second bandwidth.16 The system of claim 15, wherein the first communication link comprises an optical link, and the second communication link comprises an acoustic link.
17. The system of claim 14, wherein the transmitter transmits the latent representation and the residual via the determined available bandwidth.
18. The system of claim 11, wherein the 3D model generator is further configured to transmit the generated 3D model of the area of interest to the receiving end via a high bandwidth link.
19. The system of claim 11, wherein the 3D model generator is configured to: capture the area of interest by the camera to produce the first video sequence of the area of interest; track a movement of the camera throughout the sequence and generate accurate pose estimates for each frame; and construct the 3D model of the area of interest based on captured images and the pose estimates of the camera.
20. The system of claim 14, wherein the resolution transformer module comprises a machine learning model that dynamically adjusts the resolution of a video stream based on the available bandwidth.
21. A method for decoding data compressed by model-based video compression (MBVC), comprising: receiving a 3D model of an area of interest based on a first video sequence of the area of interest captured by a camera;receiving a latent representation of a synthesized image of the area of interest, wherein the synthesized image is generated based on the generated 3D model of the area of interest and a refined camera pose for a second video sequence of the area of interest; receiving a residual, wherein the residual is computed by comparing the synthesized image with an image captured by the camera; and reconstructing an image / video of the area of interest by incorporating the 3D model of the area of interest and the residual.
22. The method of claim 21, wherein the generated 3D model of the area of interest is received via a higher communication link, and the latent representation and the residual is received via a lower communication link.
23. A system for decoding data compressed by model-based video compression (MBVC), comprising: a model based decoder configured to: receive a 3D model of an area of interest based on a first video sequence of the area of interest captured by a camera, receive a latent representation of a synthesized image of the area of interest, wherein the synthesized image is generated based on the generated 3D model of the area of interest and a refined camera pose for a second video sequence of the area of interest, receive a residual, wherein the residual is computed by comparing the synthesized image with an image captured by the camera, and reconstruct an image / video of the area of interest by incorporating the 3D model of the area of interest and the residual.
24. The system of claim 23, wherein the model based decoder is configured to receive the generated 3D model of the area of interest via a higher communication link, and to receive the latent representation and the residual via a lower communication link.
25. A system for inspection of underwater environment, comprising,the system for MBVC and data transmission of claim 11 ; an unmanned underwater vehicle (UUV) in communication with the system for MBVC and data transmission; and a base station for establishing a wireless communication link with the UUV and a tethered link to communicate with a remote station.
26. The system of claim 25, wherein the UUV comprises a hybrid underwater vehicle (HUV).
27. The system of claim 26, wherein the HUV establishes a communication link for virtual tethering via an acoustic link for high frequency information and an optical link for short-range low latency communication.
28. The system of claim 25, wherein the UUV is equipped with the camera.
29. A system for model-based video compression (MBVC) and data transmission, comprising: one or more processors; and memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for performing or facilitating performing of the method of claim 1.
30. A system for decoding data compressed by model-based video compression (MBVC), comprising: one or more processors; and memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for performing or facilitating performing of the method of claim 21.
31. A non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors, the one or moreprograms including instructions for performing or facilitating performing of the method of claim 1 or claim 21.
Citation Information
Patent Citations
Generating video content
US11704862B1
Model-based view extrapolation for interactive virtual reality systems
US6330281B1