Vayu sense single quad depth model

The multi-aperture optical camera system with polarization information fusion addresses depth estimation challenges on reflective and transparent surfaces, enhancing accuracy and reliability in depth perception.

WO2026055271A1PCT designated stage Publication Date: 2026-03-12SENSE HOLDCO INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-09-03
Publication Date
2026-03-12

AI Technical Summary

Technical Problem

Traditional imaging systems struggle with accurate depth estimation on reflective and transparent surfaces due to their reliance on RGB wavelengths, leading to image distortions and refraction artifacts.

Method used

A system utilizing a multi-aperture optical camera system that captures polarization properties, combining RGB and polarized images through a vision transformer encoder to generate a metric depth map by fusing initial depth maps with polarization information.

Benefits of technology

Improves depth estimation accuracy, particularly on reflective and transparent surfaces, by mitigating image distortions and refraction artifacts, and enhances depth perception in real-time applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025044748_12032026_PF_FP_ABST
    Figure US2025044748_12032026_PF_FP_ABST
Patent Text Reader

Abstract

A system for improved depth estimation, the system comprising: one or more processors; and a non-transitory machine readable medium containing program instructions that are executable by the one or more processors to perform a method comprising: receiving, using one or more processors, an image; generating and outputting, using the one or more processors, an initial depth map for the image using a depth model; receiving, using the one or more processors, polarized images corresponding to the image; calculating and outputting, using the one or more processors, polarization information of the image using the polarized images and the initial depth map; and fusing the initial depth map and the polarization information to generate a metric depth map.
Need to check novelty before this filing date? Find Prior Art

Description

PCT / US25 / 44748 03 September 2025 (03.09.2025)Vayu Sense Single Quad Depth ModelCROSS REFERENCE TO RELATED APPLICATIONS

[0001] This application claims priority to and the benefit of U.S. Provisional Patent Application No. 63 / 690,733, filed September 4, 2024, and U.S. Provisional Patent Application No. 63 / 714,062, filed October 30, 2024, which are hereby incorporated by reference in their entireties and for all purposes.FIELD OF THE INVENTION

[0002] The present invention generally relates to modelling architectures and, more specifically, modelling architectures configured to optimize depth perception.BACKGROUND

[0003] Within the field of imaging, depth maps refer to images and other visual representations that contain information relating to the distance of the surfaces of objects and / or areas within an image. As such, Depth maps contain information about the distance of objects from specific perspectives, often the viewpoint of the original optical system.SUMMARY OF THE INVENTION

[0004] The purpose and advantages of the disclosed subject matter will be set forth in and apparent from the description that follows, as well as will be learned by practice of the disclosed subject matter. Additional advantages of the disclosed subject matter will be realized and attained by the methods and systems particularly pointed out in the written description and claims hereof, as well as from the appended drawings.

[0005] To achieve these and other advantages and in accordance with the purpose of the disclosed subject matter, as embodied and broadly described, the disclosed subject matter is directed to a system for improved depth estimation, the system comprising: one or more processors; and a non-transitory machine readable medium containing program instructions that are executable by the one or more processors to perform a method comprising: receiving, using one or more processors, an image; generating and outputting, using the one or more processors, an initial depth map for the image using a depth model; receiving, using the one or more processors, polarized imagesPCT / US25 / 44748 03 September 2025 (03.09.2025) corresponding to the image; calculating and outputting, using the one or more processors, polarization information of the image using the polarized images and the initial depth map; and fusing the initial depth map and the polarization information to generate a metric depth map.

[0006] In embodiments disclosed herein, the system comprises a first pass using a first vision transformer (ViT) encoder, wherein the first pass includes receiving the images and generating and outputting the initial depth map. In embodiments, the system comprises a second pass using a second vision transformer (ViT) encoder, the second pass being after the first pass, the second pass including: receiving the polarized images corresponding to the image; calculating and outputting the polarization information of the image; and fusing the initial depth map and the polarization information to generate the metric depth map. In embodiments, the first pass and the second pass are performed using a vision transformer (ViT) encoder. In embodiments, the initial depth map and the polarization information is fused within a Dense Prediction transformer (DPT). In embodiments, the polarization information comprises an Angle of Linear Polarization (AoLP) and a Degree of Linear Polarization (DoLP) of the image. In embodiments, the image is received from a multi-aperture optical camera system configured to aggregate sensor data from multiple optical sensors. In embodiments, the instructions are configured to further cause the one or more processors to: calculate the polarization information based on stereoscopic information sensed by the multi-aperture optical camera system. In embodiments, the instructions are configured to further cause the one or more processors to: calculate the polarization information by projecting each of the polarized images to a 0° polarized image space using the initial depth map. In embodiments, the image is received from a multi-aperture optical camera system configured to aggregate sensor data from multiple optical sensors. In embodiments, the instructions are configured to further cause the one or more processors to: calculate the AoLP and the DoLP based on stereoscopic information sensed by the multi-aperture optical camera system. In embodiments, the instructions are configured to further cause the one or more processors to: calculate the AoLP and the DoLP by projecting each of the polarized images to a 0° polarized image space using the initial depth map. In embodiments, each optical sensor of the multi-aperture optical camera system includesPCT / US25 / 44748 03 September 2025 (03.09.2025) a unique polarization filter applied to each aperture. In embodiments, the image is received from four optical sensors having apertures of 0°, 45°, 90°, and 135°. In embodiments, the instructions are configured to further cause the one or more processors to: calculate the AoLP and the DoLP based on stereoscopic information corresponding to the four optical sensors. In embodiments, the instructions are configured to further cause the one or more processors to: calculate the AoLP and the DoLP by projecting each of the polarized images to a 0° polarized image space using the initial depth map. In embodiments, each of the four optical sensors includes a unique polarization filter.

[0007] In accordance with another aspect of the disclosed subject matter, an improved method for generating a metric depth map of an image, the method comprising: receiving, using one or more processors, an image; generating and outputting, using the one or more processors, an initial depth map for the image using a depth model; receiving, using the one or more processors, polarized images corresponding to the image; calculating and outputting, using the one or more processors, polarization information of the image using the polarized images and the initial depth map; and fusing the initial depth map and the polarization information to generate a metric depth map.

[0008] In embodiments disclosed herein, the method comprises a first pass using a first vision transformer (ViT) encoder, wherein the first pass includes receiving the images and generating and outputting the initial depth map. In embodiments, the method comprises a second pass using a second vision transformer (ViT) encoder, the second pass being after the first pass, the second pass including: receiving the polarized images corresponding to the image; calculating and outputting the polarization information of the image; and fusing the initial depth map and the polarization information to generate the metric depth map. In embodiments, the first pass and the second pass are performed using a vision transformer (ViT) encoder. In embodiments, the initial depth map and the polarization information is fused within a Dense Prediction transformer (DPT). In embodiments, the polarization information comprises an Angle of Linear Polarization (AoLP) and a Degree of Linear Polarization (DoLP) of the image. In embodiments, the image is received from a multi-aperture optical camera system configured to aggregate sensor data from multiple optical sensors. In embodiments, calculating the polarization information based on stereoscopic information sensed by the multi-aperture opticalPCT / US25 / 44748 03 September 2025 (03.09.2025) camera system. In embodiments, calculating the polarization information by projecting each of the polarized images to a 0° polarized image space using the initial depth map. In embodiments, the image is received from a multi-aperture optical camera system configured to aggregate sensor data from multiple optical sensors. In embodiments, calculating the AoLP and the DoLP based on stereoscopic information sensed by the multi-aperture optical camera system. In embodiments, calculating the AoLP and the DoLP by projecting each of the polarized images to a 0° polarized image space using the initial depth map. In embodiments, each optical sensor of the multi-aperture optical camera system includes a unique polarization filter applied to each aperture. In embodiments, the image is received from four optical sensors having apertures of 0°, 45°, 90°, and 135°. In embodiments, calculating the AoLP and the DoLP based on stereoscopic information corresponding to the four optical sensors. In embodiments, calculating the AoLP and the DoLP by projecting each of the polarized images to a 0° polarized image space using the initial depth map. In embodiments, each of the four optical sensors includes a unique polarization filter.

[0009] In accordance with another aspect of the disclosed subject matter, one or more computer-readable non-transitory storage media including instructions that, when executed by one or more processors, are configured to cause the one or more processors to: receive, using one or more processors, an image; generate and output, using the one or more processors, an initial depth map for the image using a depth model; receive, using the one or more processors, polarized images corresponding to the image; calculate and output, using the one or more processors, polarization information of the image using the polarized images and the initial depth map; fuse the initial depth map and the polarization information to generate a metric depth map.

[0010] In accordance with embodiments disclosed herein, further including a first pass using a first vision transformer (ViT) encoder, wherein the first pass includes receiving the images and generating and outputting the initial depth map. In embodiments, further including a second pass using a second vision transformer (ViT) encoder, the second pass being after the first pass, the second pass including: receiving the polarized images corresponding to the image; calculating and outputting the polarization information of the image; and fusing the initial depth map and the polarizationPCT / US25 / 44748 03 September 2025 (03.09.2025) information to generate the metric depth map. In embodiments, the first pass and the second pass are performed using a vision transformer (ViT) encoder. In embodiments, the initial depth map and the polarization information is fused within a Dense Prediction transformer (DPT). In embodiments, the polarization information comprises an Angle of Linear Polarization (AoLP) and a Degree of Linear Polarization (DoLP) of the image. In embodiments, the image is received from a multi-aperture optical camera system configured to aggregate sensor data from multiple optical sensors. In embodiments, the instructions are configured to further cause the one or more processors to: calculate the polarization information based on stereoscopic information sensed by the multi-aperture optical camera system. In embodiments, the instructions are configured to further cause the one or more processors to: calculate the polarization information by projecting each of the polarized images to a 0° polarized image space using the initial depth map. In embodiments, the image is received from a multi-aperture optical camera system configured to aggregate sensor data from multiple optical sensors. In embodiments, the instructions are configured to further cause the one or more processors to: calculate the AoLP and the DoLP based on stereoscopic information sensed by the multi-aperture optical camera system. In embodiments, the instructions are configured to further cause the one or more processors to: calculate the AoLP and the DoLP by projecting each of the polarized images to a 0° polarized image space using the initial depth map. In embodiments, each optical sensor of the multi-aperture optical camera system includes a unique polarization filter applied to each aperture. In embodiments, the image is received from four optical sensors having apertures of 0°, 45°, 90°, and 135°. In embodiments, the instructions are configured to further cause the one or more processors to: calculate the AoLP and the DoLP based on stereoscopic information corresponding to the four optical sensors. In embodiments, the instructions are configured to further cause the one or more processors to: calculate the AoLP and the DoLP by projecting each of the polarized images to a 0° polarized image space using the initial depth map. In embodiments, each of the four optical sensors includes a unique polarization filter.PCT / US25 / 44748 03 September 2025 (03.09.2025)BRIEF DESCRIPTION OF THE DRAWINGS

[0011] The description will be more fully understood with reference to the following figures and data graphs, which are presented as exemplary embodiments of the invention and should not be construed as a complete recitation of the scope of the invention.

[0012] FIGS. 1 A - 1 B illustrate examples of optical systems used by autonomous mobile robots configured in accordance with some embodiments of the invention.

[0013] FIGS. 2A - 2B illustrate an example of a camera configured in accordance with a number of embodiments of the invention.

[0014] FIG. 3 illustrates an example of a camera integration configuration implemented in accordance with several embodiments of the invention.

[0015] FIGS. 4 - 5 conceptually illustrate modelling architectures applied to producing depth maps in accordance with many embodiments of the invention.

[0016] FIGS. 6A - 6B illustrate processes for training systems implemented in accordance with some embodiments of the invention.

[0017] FIGS. 7A - 8C illustrate images obtained using sensory mechanisms configured in accordance with various embodiments of the invention.

[0018] FIGS. 9 - 10 conceptually illustrate examples of multi-camera integration configurations implemented in accordance with certain embodiments of the invention.DETAILED DESCRIPTION

[0019] Turning now to the drawings, depth modeling systems (also referred to as depth models in this disclosure), optical systems used to capture polarized images, methods of performing polarization imaging, and methods of generating depth estimates (also referred to as depth predictions in this disclosure) in accordance with various embodiments of the invention are illustrated.

[0020] Systems configured in accordance with many embodiments may capture polarization cues, which can be used to infer depth information. Polarization images can also provide information concerning light reflections that can be relevant to operations including (but not limited to) entity detection and / or correction of poor lighting. The need for accurate and robust metric-depth estimation in real-time often clashes with the fact that (traditional) single-camera / stereo systems struggle with reflective surfaces and transparent objects (e.g., water, ice, glass, shiny materials). A major cause of this is thatPCT / US25 / 44748 03 September 2025 (03.09.2025) traditional cameras capture (only) the intensity of the incident light in the Red, Green, and Blue (RGB) wavelengths. As such, optical systems in accordance with many embodiments of the invention may use multi-aperture camera configurations to effectively capture the polarization properties of incident light. Exemplary multi-aperture camera configurations are described in U.S. Patent Publication Nos. 2025 / 0155290A1 , 2025 / 0173886A1 , and 2025 / 0227200A1 the disclosure of which are incorporated by reference herein in their entireties.

[0021] Moreover, depth modeling systems implemented in accordance with several embodiments of the invention may be configured to fuse polarization cues with depth information. By doing so, depth modeling systems configured in accordance with some embodiments of the invention can mitigate the effect of image distortions and / or refraction artifacts on depth maps. Depth modeling systems configured in accordance with many embodiments of the invention may utilize multiple different input channel modalities, corresponding to both RGB and polarized images, to fuse the latent space features in the decoder stage for more informed depth estimation. This results in improved depth estimation, in particular, for areas like reflective and transparent surfaces.

[0022] For use, in addition to or in alternative to the multi-aperture configurations and depth modeling systems described above, autonomous mobile robots, sensor systems, and modelling architectures can be utilized in configurations operating in accordance with many embodiments of the invention. Configurations may be directed to but are not limited to autonomous vehicle functionality. Autonomous vehicle functionality may include, but is not limited to architectures for updating depth estimates (and / or maps), sensory instruments, and camera configurations.

[0023] Autonomous vehicles, sensor systems that can be utilized in machine vision applications, and methods for generating high-resolution depth maps in accordance with various embodiments of the invention are discussed further below.A. Sensor Configurations

[0024] Systems and methods implemented in accordance with various embodiments of the invention may utilize image polarization and / or depth assessment for purposes including but not limited to localization, identification, and / or navigation. ManyPCT / US25 / 44748 03 September 2025 (03.09.2025) examples below may refer to “autonomous mobile robots,” however this is intended to be in a non-limiting manner. Systems in accordance with many embodiments of the invention may be used for non-vehicular autonomous robots including but not limited to security cameras. Nevertheless, examples of autonomous mobile robots implemented in accordance with some embodiments of the invention may be (e.g., unmanned) autonomous vehicles including but not limited to motor vehicles, aerial vehicles, aquatic vehicles, and / or aerospace vehicles, bipedal robots, (e.g., aerial) drones, industrial robots, autonomous robots (e.g., warehouse robots, sorting robots, factory collaboration robots, retail and hospitality robots), pet robots, robotic vacuums, pet robots, self-driving cars, healthcare related robots (e.g., hospital delivery robots, surgical assistant robots, social companion robots). Specific autonomous robot implementations may include but are not limited to one or more processors, such as a central processing unit (CPU) and / or a graphics processing unit (GPU); a data storage component; and / or one or more sensory instruments (e.g., cameras).

[0025] Examples of optical systems used by autonomous robots operating in accordance with some embodiments of the invention are illustrated in FIGS. 1A - 1 B. In many embodiments, optical systems utilizing one or more of cameras, time of flight cameras, structured illumination, light detection and ranging systems (LiDARs), laser range finders and / or proximity sensors can be utilized to acquire depth information. For example, FIG. 1 A depicts an autonomous mobile robot 100 with a sensory camera 101 A, 101 B appended to each side. FIG. 1 B depicts an autonomous mobile robot 110 topped with a wide camera rig 111 , with a pair of sensors 112A, 112B attached, one at each end of the camera rig. Nevertheless, multi-aperture configurations may follow a wide variety of camera arrangements in accordance with multiple embodiments of the invention. In many embodiments, optical systems including multi-aperture sensor (e.g., camera) configurations may be used to capture images in three dimensions that may use but are not limited to using deep stereo models. This can be applied to effectively obtain depth maps (i.e., to replicate and translate the perception of depth experienced through binocular vision). Using such configurations may result in dense per-pixel pseudo-ground truth (e.g., used for supervision). More particularly, systems and methods in accordance with some embodiments of the invention may utilize multi-aperture camera configurationsPCT / US25 / 44748 03 September 2025 (03.09.2025) for purposes including but not limited to sensing polarization properties (e.g., of incident light).

[0026] An example of a camera, operating in accordance with multiple embodiments of the invention, is illustrated in FIGS. 2A - 2B. In many embodiments, multiple image sensors can be utilized to perform depth sensing by measuring parallax, observable when images of the same scene are captured from different viewpoints / perspectives. As shown in FIG. 2A, cameras 200 configured in accordance with many embodiments may incorporate four standard image sensors 201 A, 201 B, 201 C, 201 D (e.g., in a 2x2 grid). Additionally or alternatively, each of the incorporated image sensors may have unique polarization filters applied at the corresponding apertures (enabling the capture of polarization depth cues).

[0027] In several embodiments, these collections of multiple image sensors, configured with different polarization filters, can be utilized in multi-aperture arrays to capture images of a scene at different polarization angles. Capturing images with different polarization information can enable optical systems to generate precise depth estimates using polarization cues.

[0028] As mentioned above, autonomous mobile robots may include any of a variety of sensory components for capturing or exhibiting data, including but not limited to cameras and / or sensors. In a variety of embodiments, the sensory components can be used to gather inputs and / or provide outputs, either of which can be used to localize and / or navigate robots. Additional sensors may include but are not limited to ultrasonic sensors, motion sensors, light sensors, infrared sensors, and / or custom sensors.

[0029] Systems in accordance with numerous embodiments of the invention may allow for the use of polarized images alongside (or in alternative to) standard RGB images (e.g., for determining / improving depth perception). Polarization information used for calculations including but not limited to depth estimation can be represented as (at least) two additional channels of information at each pixel, including but not limited to: Angle of Linear Polarization and Degree of Linear Polarization. The Angle of Linear Polarization (AoLP) refers to the angle of maximum polarization, while the Degree of Linear Polarization (DoLP) refers to the ratio of -- - - - - - - — . Both can be intensity of the un polarized part of the light effectively computed from the captured polarized images in a straightforward way whenPCT / US25 / 44748 03 September 2025 (03.09.2025) camera images are taken from the same viewpoint (e.g., using a camera with a single aperture and an image sensor that has polarization filters applied at the pixel level), but tend to be more difficult to measure for multi-aperture camera configurations. Therefore, system architectures implemented in accordance with a number of embodiments of the invention (and described below) may allow for polarized images to be captured from slightly displaced apertures (i.e. , no longer aligned at the pixel level), while still allowing for derivation of the AoLP and DoLP.

[0030] Hardware-based central computers (e.g., processors) may be implemented within autonomous mobile robots and other devices operating in accordance with various embodiments of the invention to execute program instructions and / or software, causing computers to perform various methods and / or tasks, including the techniques described herein. Several functions including but not limited to data processing, data collection, machine learning operations, and simulation generation can be implemented on singular processors, on multiple cores of singular computers, and / or distributed across multiple processors. Central computers may take various forms including but not limited to CPUs, digital signal processors (DSP), core processors within Application Specific Integrated Circuits (ASIC), and / or GPUs for the manipulation of computer graphics and image processing, as described below. CPUs may be directed to autonomous mobile robot operations including (but not limited to) the application of sensory input, polarization of obtained images, and / or operation of depth modeling systems. As indicated above, CPUs may, additionally or alternatively, be coupled with one or more GPUs. GPUs may be directed towards, but are not limited to ongoing perception and sensory efforts.

[0031] Central computers implemented in accordance with numerous embodiments of the invention may be configured to process input data according to instructions stored in data storage components. Data storage components may include but are not limited to hard disk drives, nonvolatile memory, and / or other non-transient storage devices. Data storage components, including but not limited to memory, can be loaded with software code that is executable by processors to achieve certain functions. Memory may exist in the form of tangible, non-transitory, computer-readable mediums configured to store instructions that are executable by processors. Data storagePCT / US25 / 44748 03 September 2025 (03.09.2025) components may be further configured to store supplementary information including but not limited to sensory data and / or depth estimates.

[0032] In some embodiments of the invention, central computers may be coupled to at least one interface component including but not limited to Jetson AGX Orin camera expansion boards and / or Display Serial Interfaces following Mobile Industry Processor Interface Protocols (MIPI DSIs). Camera expansion boards in accordance with certain embodiments of the invention may allow for the combination of outputs from a plurality of cameras into individual serial links including but not limited to Gigabit Multimedia Serial Links (e.g., GMSL, GMSL2, GMSL3). GMSL configurations have functionality elaborated on GMSL2 Channel Specification User Guide, Analog Devices, Inc., https: / / www.analog.com / media / en / technical-documentation / user-guides / gmsl2-channel- specification-user-guide.pdf / , the entire disclosure of which, including the disclosure related to camera integration, is hereby incorporated by reference in its entirety.

[0033] A serial link configuration applied in accordance with numerous embodiments of the invention is illustrated in FIG. 3. As indicated above, serial link 340 configurations may be used to transmit sensory data 335 from cameras 310, including but not limited to quad cameras, to central computers 345, including but not limited to Nvidia Jetson AGX Orin central (“Orin”) computing units. In accordance with multiple embodiments of the invention, the same serial links 340 used to connect the cameras 310 to the central computers 345 may additionally (or alternatively) carry power required for the cameras 310. In such cases, the required power may be carried through the use of power over coax (POC) cables. In accordance with numerous embodiments, the POC cables may have various lengths (e.g., 3 meters or less). Additionally or alternatively, power for the POC can be drawn from (e.g., Orin) central computer itself and / or from different external DC source(s).

[0034] In accordance with multiple embodiments, within the cameras 310, sensory data from one or more associated sensors 330 can undergo a first level of aggregation and / or virtual channel creation. Cameras 310 configured in accordance with many embodiments of the invention may aggregate sensor data 335 from multiple (e.g., 4) sensors 330, where the sensor data 335 can be RAW and / or processed (e.g., by ISPs). Systems configured in accordance with some embodiments of the invention may usePCT / US25 / 44748 03 September 2025 (03.09.2025)Image Signal Processors (ISPs) to transform RAW images from the sensor(s) into high- quality image(s) for machine vision and tele-assist (to display the surround vision user). This may be used to save computational power and / or memory (e.g., dynamic randomaccess memory) of the central computer(s) and / or for the software algorithms used by the system. Additionally or alternatively, this may allow for quick isolation of problems to computer and / or sensor boundaries.

[0035] In accordance with some embodiments of the invention, the sensor data 335 may be aggregated prior to reaching the serial links 340 using various aggregators including but not limited to MIPI aggregators 325 (e.g., Lattice CrossLink-NX MIPI aggregators). Additionally or alternatively, within the cameras 310, singular MIPI buses may be used to transmit the (aggregated) sensor data to the serial links 340 via serializer. When GMSL is used alongside Orin computing units, each of the individual GMSL links may be able to accommodate up to sixteen virtual channels / streams. Further, as indicated in FIGS. 3B and 3C, in accordance with various embodiments of the invention, as many as eight different cameras 310 may be connected to the central computer 345 (e.g., using GMSL links 340).

[0036] In accordance with multiple embodiments of the invention, the serial links 340 may be enabled through the use of components including but not limited to serializers (corresponding to the camera(s) 310), deserializers (corresponding to the central computer(s)), and / or connectors (corresponding to both). In accordance with some embodiments, serial links 340 may use easily accessible connectors (e.g., Fakra connectors). In such cases, the (e.g., Fakra) connectors can be quad, dual, and / or single based on the logistics and mechanical constraints.

[0037] Interfaces used by systems configured in accordance with multiple embodiments of the invention may (additionally or alternatively) take the form of one or more wireless interfaces and / or one or more wired interfaces. In accordance with many embodiments of the invention, interfaces may be used to communicate with other devices and / or components including but not limited to the image sensors of cameras, as described above. Moreover, additional interfaces configured in accordance with many embodiments of the invention can be used to integrate imaging systems into greaterPCT / US25 / 44748 03 September 2025 (03.09.2025) configurations that may be applied to depth modeling (e.g., in autonomous mobile robots) as described below.

[0038] While specific autonomous mobile robots and optical systems are described above with reference to FIGS. 1A - 3, any of a variety of sensory configurations can be implemented as appropriate to the requirements of specific applications in accordance with some embodiments of the invention. Furthermore, applications and methods in accordance with various embodiments of the invention are not limited to use within any specific autonomous systems. Accordingly, it should be appreciated that the implementations described herein can also be implemented outside the context of an autonomous mobile robot described above with reference to FIGS. 1 A - 3. Many systems and methods for implementing depth modeling and applications in accordance with numerous embodiments of the invention are discussed further below.B. Modeling Configurations

[0039] As noted above, autonomous mobile robots implemented in accordance with many embodiments of the invention may utilize depth modeling systems in order to perform functions associated with localization. In many instances, the depth modeling systems utilize inputs from sensor systems (as described above) that contain depth cues and / or polarization cues.

[0040] In doing so, depth modelling systems implemented in accordance with several embodiments of the invention may combine polarization and depth cues in order to refine depth estimates. Systems may do so by integrating polarization cues / features including but not limited to AoLP estimates, DoLP estimates, and / or preliminary (i.e., approximate) depth estimates with three-channel (e.g., regular RGB-based) transformer architectures. In many embodiments of the invention, these transformer architectures may be designed to process information from different apertures, ultimately improving the reliability of depth estimation across various scenarios and environments.

[0041] In many embodiments of the invention, the displacement that results from combining multiple camera apertures can lead to image distortion in the resulting outputs (e.g., depth maps). However, this feature, additionally or alternatively, carries supplementary stereoscopic information. For example, the differing aperture positions mean that the amount of shift (between the same feature(s) across cameras) canPCT / US25 / 44748 03 September 2025 (03.09.2025) effectively provide evidence about the depth of the feature(s) in the area. The disclosed system architecture is designed to make use of this extra information to solve the above alignment problem.

[0042] Certain implementations of the invention may operate as variations on the Depth Anything V2 system, a transformer-based architecture that leverages Vision Transformers (ViTs) as the encoders and a specialized DPT head as the decoder. This depth model is designed to predict depth from RGB images, and features several critical components. The proposed model (“Depth Anything V2 with polarization) is different from the base architecture in (1 ) having a separate ViT encoder for the polarization channels, and (2) in the feature fusion layers, including polarization feature embeddings.

[0043] Conceptual illustrations of system architectures applied as depth modelling systems in accordance with some embodiments of the invention are depicted in FIGS. 4 - 5. As mentioned above, system architectures (e.g., for depth modelling systems) configured in accordance with many embodiments of the invention may be built upon the effective integration of polarization information (e.g., four-channel polarization features 420) with traditional visual cues (e.g., RGB 410 images). System architectures may include but are not limited to two (or more) encoders 430, 440 and one (or more) depth decoder 450. The system architecture disclosed in FIG. 4 contains two Vision Transformer (ViT) encoders 430, 440, and one Dense Prediction Transformer (DPT) head as the depth decoder 440.

[0044] System architectures implemented in accordance with multiple embodiments of the invention may start by fine-tuning the encoder 430 to decoder 450 pipeline that estimates depth using standard visual features extracted from RGB 410 channels. For this, systems implemented in accordance with several embodiments of the invention may pre-trained weights (e.g., Depth Anything V2 pre-trained weights) for the encoder 430. Additionally or alternatively, systems may fine-tune the decoder(s) 450 with domain-specific data.

[0045] Furthermore, using these pipelines given RGB 410 images, systems can obtain approximate depth estimates. These (initial) approximate depth estimates may serve as a guiding baseline for the depth modeling systems. Systems may, additionally or alternatively, compute polarization features 420 including but not limited to the AnglePCT / US25 / 44748 03 September 2025 (03.09.2025) of Linear Polarization (AoLP) and Degree of Linear Polarization (DoLP) from predicted approximate depth estimates and / or obtained polarized images. The obtained polarized images used for these computations may include but are not limited to those obtained from the optical system (for example, four polarized images from apertures at 0°, 45°, 90°, and 135°). In accordance with multiple embodiments of the invention, systems may compute the polarization information including but not limited to AoLP and DoLP values by projecting all the polarized images to 0° polarized image space using the approximate depth estimates, thereby addressing the pixel space misalignment of the polarized images that arises due to the sensor’s displaced apertures.

[0046] Systems may, additionally or alternatively, concatenate initial depth maps (e.g., determined from the approximate depth estimates) with polarization 420 channels, generated based on the above polarization information, to begin a pass through the second encoder 440 to decoder 450 pipeline. This pipeline may take derived polarization 420 channels including but not limited to the approximate depth, the cosine and / or sine of the AoLP, and the DoLP as inputs.

[0047] As discussed above, the output features from the two encoders 430, 440 may be fused within the decoder(s) 450 in order to refine the depth estimates. In doing so, certain decoders 450 (e.g., the DPT decoderfrom Depth Anything V2) can be modified to have twice the number of input channels to incorporate this fusion. The decoder(s) 450 use this fused embedding to output corrected metric depth 460 maps.

[0048] In essence, depth models implemented in accordance with many embodiments of the invention may utilize two different input channel modalities (three- channel and four-channel) through two ViT encoders and fuse the latent space features in the decoder stage for more informed depth estimation. The result is a more refined depth map that benefits from the additional polarization cues, particularly in areas where traditional depth estimation methods struggle, such as reflective and transparent surfaces. This particular approach has several major benefits:

[0049] (1 ) Preservation of Pretrained Efficacy: By maintaining the integrity of theDepth Anything V2 model’s encoder, system architectures ensure that the model’s ability to extract and process traditional RGB features is not diminished. The decoder-stagePCT / US25 / 44748 03 September 2025 (03.09.2025) fusion of polarization cues allows the model to refine its depth predictions based on additional information without retraining the entire network;

[0050] (2) Interpretable Polarization Correction: The system architecture is designed to be interpretable by incorporating the approximate depth as an input to a polarization “correction mechanism.” This allows the system architecture to explicitly measure the effect of the polarization cues on the final depth prediction. As the network only begins learning with the introduction of polarization data at the decoder stage, systems implemented in accordance with some embodiments of the invention can isolate and quantify the improvements brought about by the polarization correction; and

[0051] (3) Scalable and Efficient Integration: The modular nature of the system architecture enables easy scaling and adaptation to different datasets and environments. By focusing on the decoder-stage learning, system architectures implemented in accordance with various embodiments minimize the computational overhead while maximizing the benefits of the polarization cues.

[0052] FIG. 5 illustrates a (fine-tuned) variation on the same depth modelling system architecture disclosed in FIG. 4, operating in accordance with several embodiments of the invention. As shown in FIG. 5, post-finetuning, the depth modelling system demonstrates remarkable sim-to-real transfer capabilities, achieving an error as low as 3% on the in-domain validation set. In accordance with some embodiments of the invention, a single inference (e.g., depth map) may require two forward passes: the first pass can produce approximate depth and the second can produce refined depth estimates. Using the Depth Anything V2 system, passes may take 40ms (median), allowing a total inference of roughly 80ms. This enables the model to run at 10 frames per second comfortably.

[0053] In accordance with many embodiments of the invention, some scenarios may require training depth modeling systems without the availability of objective benchmarks (e.g., ground truth depth values) corresponding to the data (e.g., features) captured from the optical systems. To overcome this, systems in accordance with various embodiments of the invention may employ a hybrid training strategy. Systems can utilize polarization-simulating simulators (e.g., Mitsuba rendering frameworks) to generate synthetic polarized data. By leveraging these advanced simulation capabilities, systemsPCT / US25 / 44748 03 September 2025 (03.09.2025) in accordance with many embodiments of the invention can mimic real-world lighting and material interactions. However, recognizing the limited availability of real-world polarized data, in accordance with certain embodiments of the invention, primary depth modelling systems can be trained on the (far more abundant) conventional RGB images. When the primary depth modelling systems are trained, polarization channels can be introduced as a corrective mechanism during fine-tuning.

[0054] The above strategy ensures that the core visual feature extraction capabilities of depth modelling systems configured in accordance with various embodiments of the invention can remain intact, while the additional polarization information may refine and correct depth predictions, particularly in challenging scenarios.

[0055] To further mitigate the data scarcity issue, systems and methods in accordance with miscellaneous embodiments of the invention may incorporate lighting and viewpoint-based AoLP / DoLP signatures derived from (e.g., unpolarized, four- aperture) optical systems. By utilizing these signatures, systems can approximate the polarization effects in simulators / situations where actual polarized data is unavailable. This method enhances the robustness of the model without the need for extensive polarized datasets, making it both practical and scalable.

[0056] A two-phase process for training systems implemented in accordance with some embodiments of the invention is illustrated in FIGS. 6A - 6B. Systems including but not limited to the depth modeling systems described above may be trained using the process phases described above.

[0057] FIG. 6A illustrates Phase 1 of the process, where a first depth modeling system (e.g., a metric Depth Anything V2 model) may be used to produce approximate metric depth from (e.g., singular) images. As such, the goal of Phase 1 is to train a model that can quickly produce effective initial depth estimates (e.g., sufficient for approximating information including but not limited to AoLP and DoLP).

[0058] Process 600 evaluates (605) the performance of a plurality of depth models to select an optimal model. In accordance with many embodiments of the invention, high performance may be interpreted relative to a variety of factors including but not limited to real-time inference capability. As an example, previous evaluations have been performed on the inference capability of three core Depth Anything v2 models (small-core, base-PCT / US25 / 44748 03 September 2025 (03.09.2025) core, and large-core model) on an NVIDIA Jetson Orin platform, while model checkpoints were exported to Open Neural Network Exchange (ONNX) and TensorRT formats to facilitate on-device evaluation. The small-core model was selected due to its potential to meet a pre-determined 10 Hz inference target, which is critical for real-time applications.

[0059] Process 600 collects (610) data from a plurality of simulators to create a training set. In accordance with many embodiments of the invention, synthetic and / or authentic data may be collected for multiple domains (e.g., outdoor, indoor) and / or scenarios to ensure that the training set is sufficiently diverse. Additionally or alternatively, in accordance with several embodiments of the invention, training images may be resized to a consistent resolution. A previous dataset included approximately 2 million images that were resized to a square 518x518 resolution. This resolution choice was driven by the need to meet the multiple-of-14 constraint imposed by the DinoV2 encoder and to align with the Depth Anything v2 fine-tuning methodology.

[0060] Process 600 conducts (615) training of the optimal model based on the diverse training set. To replicate the strong training regimen of Depth Anything v2, past trainings have used the original iteration step-based learning rate settings and scheduler. This allowed for maintaining the integrity of the training dynamics while finetuning. Previous trainings were also conducted on a dedicated NVIDIA RTX A6000 GPU, with a batch size of 48, over 300,000 steps. This setup provided the computational power necessary to fine-tune the small-core Depth Anything v2 model effectively.

[0061] In accordance with several embodiments of the invention, process 600 may implement (620) a maximum, depth-clamping strategy for depth estimates. For example, implementing (620) a maximum, depth-clamping strategy may be performed to manage the presence of pixels that correspond to infinite depth (e.g., sky pixels). This approach can cap all pixels at a pre-specified maximum depth value, enabling the models configured in accordance with several embodiments of the invention to penalize incorrect predictions even when they fall outside the expected range.

[0062] FIG. 6B illustrates Phase 2 where a second “polarization correction” depth modeling system (e.g., a polarization-aware metric depth model) may be trained to refine depth measurements by inputting the approximate metric depth produced in Phase 1 . As suggested above, in accordance with some embodiments of the invention, both phasesPCT / US25 / 44748 03 September 2025 (03.09.2025) may be used to train the same model (e.g., a Depth Anything V2 model with Polarization). Process 650 collects (655) polarized image data and ground truth depth data from a plurality of sources. In accordance with some embodiments of the invention, potential sources of data have included but are not limited to:

[0063] (1 ) Synthetic data from polarization-simulating simulators (e.g., MitsubaTenderer): Mitsuba can simulate the polarization properties of light taking the light source and types of materials in the scene into account. This allows for the creation of highly diverse scenes with nearly perfect polarized images which are nearly impossible to generate in the real world;

[0064] (2) Synthetic data from alternative, non-polarization-simulating, simulators(e.g., Unreal Engine and / or Nvidia Omniverse): Synthetic data from alternative simulators, including but not limited to the above, may be used to contribute diversity to the data eventually used to train the depth modeling system. Additionally or alternatively, nonpolarization-simulating data may be used for simulating non-polarization-related visual effects that occur with optical systems configured in accordance with several embodiments of the invention; and / or

[0065] (3) Real data (e.g., from baseline stereo rigs): Past implementations have used methods including but not limited to wide-baseline stereo rigs (e.g., the rig disclosed in FIG. 1 B) for sources of non-synthetic data. The wide baseline rigs have been used to produce depth estimates using strong off-the-shelf deep stereo models. This results in dense per-pixel pseudo-ground truth which may be used for supervision.

[0066] Process 650 trains (660) depth modeling system based on the polarized image data and ground truth depth. In accordance with multiple embodiments of the invention, several portions of the depth modeling system (including but not limited to the RGB portion(s)) may be initialized (e.g., using Phase 1 models) and kept frozen. Additionally or alternatively, at least some of the depth modeling system may be trained using random initialization.

[0067] Process 650 inputs (665) set of initial approximate depth estimates into the system. The above arrangement can allow for the retention of general depth estimation capabilities while allowing the depth modeling system to learn to refine its predictionsPCT / US25 / 44748 03 September 2025 (03.09.2025) using data including but not limited to polarization and small baseline stereoscopic cues. Process 650 obtains (670), from the depth modeling system, refined depth estimates.

[0068] While specific processes for training depth modeling systems are described above in FIGS. 6A - 6B, any of a variety of processes can be utilized to train systems as appropriate to the requirements of specific applications. In certain embodiments, steps may be executed or performed in any order or sequence not limited to the order and sequence shown and described. In a number of embodiments, some of the above steps may be executed or performed substantially simultaneously where appropriate or in parallel to reduce latency and processing times. In some embodiments, one or more of the above steps may be omitted.

[0069] While specific modeling system configurations are described above with reference to FIGS. 4 - 6B, any of a variety of approaches to depth modeling can be implemented in accordance with multiple embodiments of the invention. Furthermore, applications and methods in accordance with various embodiments of the invention are not limited to use within any specific autonomous systems. Accordingly, it should be appreciated that the implementations described herein can also be implemented outside the context of an autonomous mobile robot described above.C. Imaging Configurations

[0070] Polarization imaging configurations may be applied to purposes including but not limited to distinguishing road hazards, as is evident in FIGS. 7A - 8C. Polarization information can be especially helpful in enabling machine vision applications to modify sensor output. In such cases, polarization sensors may be especially effective when used in analyzing high-dynamic range scenes. Referring now to the parking lot depicted in FIG. 7A, the challenges of interpreting an image of a high dynamic scene containing objects that are in shadow can be appreciated. Specifically, many potential hazards are difficult to identify which, in fast-paced circumstances, can quickly and easily lead to accidents. In contrast, FIG. 7B illustrates an example of an image generated using a polarization imaging system configured in accordance with several embodiments of the invention. This figure emphasizes how clearly objects can be discerned in high dynamic range images using polarization information. Using such systems, various dangers (such as the variousPCT / US25 / 44748 03 September 2025 (03.09.2025) hidden objects revealed in FIG. 7B) may be converted to a much easier-to-navigate format.

[0071] Systems and methods operating in accordance with certain embodiments of the invention may maximize effectiveness at distinguishing road impediments and / or minimize the danger associated with road distractions. Potential road impediments localized by sensors implemented in accordance with some embodiments of the invention may include but not limited to ice, oil spills, and potholes. The following pair of images illustrate an example of black ice on a partly frozen road being made far more visible by converting the input image of FIG. 7C to an (RGB) polarized image in FIG. 7D. Additionally or alternatively, potential distractions mitigated by sensors may include but are not limited to reflections from standing water, snow, and sun glare.

[0072] FIGS. 7E - 7G illustrate depth models implemented, in accordance with multiple embodiments of the invention, to incorporate polarization information including but not limited to DoLP and AoLP values. Past evaluations have determined that polarization information was instrumental in correcting depth errors related to reflective and shiny surfaces. By fusing polarization cues with depth information, polarization-aware Depth Anything V2 models implemented in accordance with several embodiments of the invention can mitigate the effect of image distortions and / or refraction artifacts. In some implementations, refraction artifacts like ghosting and duplication have been minimized.

[0073] As suggested above, training data from a plurality of sources may be used for training depth modeling systems in accordance with many embodiments of the invention. Training data used in accordance with many embodiments of the invention may include but are not limited to image data and / or ground truth depth data. FIG. 7E discloses four examples of depth maps (right) generated from specific image inputs (left) and training data. Moreover, the image data and / or ground truth depth data used for the four examples were produced by (from top to bottom): (1 ) a Mitsuba Tenderer, (2) an Unreal engine, (3) a Nvidia Omniverse platform, and (4) a wide baseline rig. FIGS. 7F and 7G depict additional depth models produced by polarization-aware depth modeling systems implemented in accordance with several embodiments of the invention. In particular, the increase in prediction effectiveness can be seen for both outdoor (FIG. 7F) and indoor (FIG. 7G) examples.PCT / US25 / 44748 03 September 2025 (03.09.2025)

[0074] As indicated above, images generated by systems operating in accordance with various embodiments of the invention may be configured using various polarizations and / or (combinations of) channels including but not limited to greyscale and (one or more) RGB channels. In accordance with some embodiments of the invention, channel 1 may be used to optimize the identification of material properties. Additionally or alternatively, channel 2 may be used to distinguish object shape and / or surface texture. This is reflected in the parking lot depicted in FIG. 8A, wherein the respective impacts of oil and water on a dry surface are emphasized. In channel 1 , the brightness that the water and (especially) the oil exhibit compared to the dry surface is fairly pronounced, thereby distinguishing the two materials. Meanwhile, in channel 2, the (liquid) surface texture makes the water- covered and oil-covered areas look fairly consistent with the dry (yet flat) surface, albeit with their respective boundaries evident. As shown in the two roadside examples disclosed in FIGS. 8B and 8C, images generated using various polarizations may (additionally or alternatively) result in images with dampened reflections (e.g. , in channel 1 ) and / or distinguishable areas with more distinct surface textures (e.g., the identifiable snow and ice in channel 2).

[0075] Finally, as suggested above, depth map generation may vary substantially according to camera configuration. Examples of multi-camera integration configurations implemented in accordance with multiple embodiments of the invention are illustrated in FIG. 9 - 10. In many implementations, as many as eight different cameras may be connected to given central computers (e.g., using serial links). In such instances, camera expansion boards may have (e.g., GMSL) quad deserializers, with four serial link inputs and / or two MIPI Buses (as depicted in FIG. 9). In such cases, for quad deserializer-based design, two inputs of the serial link (e.g., GMSL) can share singular MIPI buses. As in FIG. 9, four serial link inputs can use the two MIPI buses of a central computer. Additionally or alternatively, as in FIG. 10, systems in accordance with various embodiments may have the central computer be connected to eight different serial link inputs / streams. To enable this, systems may have two (or more) GMSL quad deserializers working in tandem with the central computer. Additionally or alternatively, in accordance with some embodiments, a pair of quad cameras may be used for stereo vision alongside singular sensor-based cameras to minimize power usage and / orPCT / US25 / 44748 03 September 2025 (03.09.2025) optimize cost. Moreover, when IMUs are incorporated, systems configured in accordance with certain embodiments may include three-axis gyroscopes to periodically confirm camera pose(s). Based on the IMUs, systems may synchronize the capture with video capture, may match the IMU axis to the camera axis, and / or may make the Z-axis the optical path.

[0076] While specific applications are described above for utilizing polarized images to generate depth models, any of a variety of processes can be utilized as appropriate to the requirements of specific applications in accordance with some embodiments of the invention. Furthermore, systems and methods in accordance with multiple embodiments of the invention are not limited to use within vehicular and / or navigational systems. Accordingly, it should be appreciated that the systems described herein can also be implemented outside the environmental contexts described above with reference to FIGS. 7A - 9.

[0077] Embodiments disclosed herein include:

[0078] Aspects of the disclosed subject matter are set out in the claims and additional features are further described in the dependent claims and within the disclosure. The additional features may be provided in combination within a system of method as disclosed herein. Moreover, a system or method in accordance with the disclosed subject matter may be provided that combines features of the independent claims together with any of the features of the dependent claims.

[0079] A. A system for improved depth estimation, the system comprising: one or more processors; and a non-transitory machine readable medium containing program instructions that are executable by the one or more processors to perform a method comprising: receiving, using one or more processors, an image; generating and outputting, using the one or more processors, an initial depth map for the image using a depth model; receiving, using the one or more processors, polarized images corresponding to the image; calculating and outputting, using the one or more processors, polarization information of the image using the polarized images and the initial depth map; and fusing the initial depth map and the polarization information to generate a metric depth map.PCT / US25 / 44748 03 September 2025 (03.09.2025)

[0080] B. An improved method for generating a metric depth map of an image, the method comprising: receiving, using one or more processors, an image; generating and outputting, using the one or more processors, an initial depth map for the image using a depth model; receiving, using the one or more processors, polarized images corresponding to the image; calculating and outputting, using the one or more processors, polarization information of the image using the polarized images and the initial depth map; and fusing the initial depth map and the polarization information to generate a metric depth map.

[0081] C. One or more computer-readable non-transitory storage media including instructions that, when executed by one or more processors, are configured to cause the one or more processors to: receive, using one or more processors, an image; generate and output, using the one or more processors, an initial depth map for the image using a depth model; receive, using the one or more processors, polarized images corresponding to the image; calculate and output, using the one or more processors, polarization information of the image using the polarized images and the initial depth map; fuse the initial depth map and the polarization information to generate a metric depth map.

[0082] Each of embodiment A, B, and C may have one or more of the following additional elements in any combination:

[0083] Element 1 : wherein the method comprises a first pass using a first vision transformer (ViT) encoder, wherein the first pass includes receiving the images and generating and outputting the initial depth map.

[0084] Element 2: wherein the method comprises a second pass using a second vision transformer (ViT) encoder, the second pass being after the first pass, the second pass including: receiving the polarized images corresponding to the image; calculating and outputting the polarization information of the image; and fusing the initial depth map and the polarization information to generate the metric depth map.

[0085] Element 3: wherein the first pass and the second pass are performed using a vision transformer (ViT) encoder.

[0086] Element 4: wherein the initial depth map and the polarization information is fused within a Dense Prediction transformer (DPT).PCT / US25 / 44748 03 September 2025 (03.09.2025)

[0087] Element 5: wherein the polarization information comprises an Angle of Linear Polarization (AoLP) and a Degree of Linear Polarization (DoLP) of the image.

[0088] Element 6: wherein the image is received from a multi-aperture optical camera system configured to aggregate sensor data from multiple optical sensors.

[0089] Element 7: wherein the instructions are configured to further cause the one or more processors to: calculate the polarization information based on stereoscopic information sensed by the multi-aperture optical camera system.

[0090] Element 8: wherein the instructions are configured to further cause the one or more processors to: calculate the polarization information by projecting each of the polarized images to a 0° polarized image space using the initial depth map.

[0091] Element 9: wherein the image is received from a multi-aperture optical camera system configured to aggregate sensor data from multiple optical sensors.

[0092] Element 10: wherein the instructions are configured to further cause the one or more processors to: calculate the AoLP and the DoLP based on stereoscopic information sensed by the multi-aperture optical camera system.

[0093] Element 11 : wherein the instructions are configured to further cause the one or more processors to: calculate the AoLP and the DoLP by projecting each of the polarized images to a 0° polarized image space using the initial depth map.

[0094] Element 12: wherein each optical sensor of the multi-aperture optical camera system includes a unique polarization filter applied to each aperture.

[0095] Element 13: wherein the image is received from four optical sensors having apertures of 0°, 45°, 90°, and 135°.

[0096] Element 14: wherein the instructions are configured to further cause the one or more processors to: calculate the AoLP and the DoLP based on stereoscopic information corresponding to the four optical sensors.

[0097] Element 15: wherein the instructions are configured to further cause the one or more processors to: calculate the AoLP and the DoLP by projecting each of the polarized images to a 0° polarized image space using the initial depth map.

[0098] Element 16: wherein each of the four optical sensors includes a unique polarization filter.PCT / US25 / 44748 03 September 2025 (03.09.2025)

[0099] Element 17: comprising a first pass using a first vision transformer (ViT) encoder, wherein the first pass includes receiving the images and generating and outputting the initial depth map.

[0100] Element 18: comprising a second pass using a second vision transformer (ViT) encoder, the second pass being after the first pass, the second pass including: receiving the polarized images corresponding to the image; calculating and outputting the polarization information of the image; and fusing the initial depth map and the polarization information to generate the metric depth map.

[0101] Element 19: wherein the first pass and the second pass are performed using a vision transformer (ViT) encoder.

[0102] Element 20: wherein the initial depth map and the polarization information is fused within a Dense Prediction transformer (DPT).

[0103] Element 21 : wherein the polarization information comprises an Angle of Linear Polarization (AoLP) and a Degree of Linear Polarization (DoLP) of the image.

[0104] Element 22: wherein the image is received from a multi-aperture optical camera system configured to aggregate sensor data from multiple optical sensors.

[0105] Element 23: wherein the polarization information comprises an Angle of Linear Polarization (AoLP) and a Degree of Linear Polarization (DoLP) of the image.

[0106] Element 24: wherein the image is received from a multi-aperture optical camera system configured to aggregate sensor data from multiple optical sensors.

[0107] Element 25: further comprising calculating the polarization information based on stereoscopic information sensed by the multi-aperture optical camera system.

[0108] Element 26: further comprising calculating the polarization information by projecting each of the polarized images to a 0° polarized image space using the initial depth map.

[0109] Element 27: wherein the image is received from a multi-aperture optical camera system configured to aggregate sensor data from multiple optical sensors.

[0110] Element 28: wherein the Angle of Linear Polarization (AoLP) and Degree of Linear Polarization (DoLP) is calculated based on stereoscopic information sensed by the multi-aperture optical camera system.PCT / US25 / 44748 03 September 2025 (03.09.2025)

[0111] Element 29: wherein the AoLP and the DoLP is calculated by projecting each of the polarized images to a 0° polarized image space using the initial depth map.

[0112] Element 30: wherein each optical sensor of the multi-aperture optical camera system includes a unique polarization filter.

[0113] Element 31 : wherein the image is received from optical sensors having apertures of 0°, 45°, 90°, and 135°.

[0114] Element 32: wherein the AoLP and the DoLP is calculated based on stereoscopic information corresponding to the four optical sensors.

[0115] Element 33: wherein the AoLP and the DoLP is calculated by projecting each of the polarized images to a 0° polarized image space using the initial depth map.

[0116] Element 34: wherein each of the four optical sensors includes a unique polarization filter.

[0117] Element 35: further comprising a first pass using a first vision transformer (ViT) encoder, wherein the first pass includes receiving the images and generating and outputting the initial depth map.

[0118] Element 36: further comprising a second pass using a second vision transformer (ViT) encoder, the second pass being after the first pass, the second pass including: receiving the polarized images corresponding to the image; calculating and outputting the polarization information of the image; and fusing the initial depth map and the polarization information to generate the metric depth map.

[0119] Element 37 : wherein the first pass and the second pass are performed using a vision transformer (ViT) encoder.

[0120] Element 38: wherein the initial depth map and the polarization information is fused within a Dense Prediction transformer (DPT).

[0121] Element 39: wherein the polarization information comprises an Angle of Linear Polarization (AoLP) and a Degree of Linear Polarization (DoLP) of the image.

[0122] Element 16: wherein the image is received from a multi-aperture optical camera system configured to aggregate sensor data from multiple optical sensors.

[0123] Element 40: wherein the instructions are configured to further cause the one or more processors to: calculate the polarization information based on stereoscopic information sensed by the multi-aperture optical camera system.PCT / US25 / 44748 03 September 2025 (03.09.2025)

[0124] Element 41 : wherein the instructions are configured to further cause the one or more processors to: calculate the polarization information by projecting each of the polarized images to a 0° polarized image space using the initial depth map.

[0125] Element 42: wherein the image is received from a multi-aperture optical camera system configured to aggregate sensor data from multiple optical sensors.

[0126] Element 43: wherein the instructions are configured to further cause the one or more processors to: calculate the AoLP and the DoLP based on stereoscopic information sensed by the multi-aperture optical camera system.

[0127] Element 44: wherein the instructions are configured to further cause the one or more processors to: calculate the AoLP and the DoLP by projecting each of the polarized images to a 0° polarized image space using the initial depth map.

[0128] Element 45: wherein each optical sensor of the multi-aperture optical camera system includes a unique polarization filter applied to each aperture.

[0129] Element 46: wherein the image is received from four optical sensors having apertures of 0°, 45°, 90°, and 135°.

[0130] Element 47: wherein the instructions are configured to further cause the one or more processors to: calculate the AoLP and the DoLP based on stereoscopic information corresponding to the four optical sensors.

[0131] Element 48: wherein the instructions are configured to further cause the one or more processors to: calculate the AoLP and the DoLP by projecting each of the polarized images to a 0° polarized image space using the initial depth map.

[0132] Element 49: wherein each of the four optical sensors includes a unique polarization filter.

[0133] D. An improved method for training a depth estimation model, the method comprising: receiving, using one or more processors, an image data set from a plurality of simulators; training, using the one or more processors, a first depth model using the image data set; generating, using the one or more processors, a set of initial depth estimates using the first depth model; receiving, using the one or more processors, polarized image data and ground truth depth from a plurality of data sources; training, using the one or more processors, a second depth model using the received polarized image data and the ground truth data; inputting, using the one or more processors, thePCT / US25 / 44748 03 September 2025 (03.09.2025) set of initial depth estimates into the second depth model; refining, using the one or more processors, the initial depth estimates based on the training and inputting; and outputting, using the one or more processors, refined depth estimates.

[0134] E. One or more computer-readable non-transitory storage media including instructions that, when executed by one or more processors, are configured to cause the one or more processors to: in a first phase using a first vision transformer (ViT) encoder and a Dense Prediction transformer (DPT) decoder: receive, from a plurality of simulators, an image data set; train a first depth model using the image data set; generate a set of initial depth estimates using the first depth model; in a second phase using a second ViT encoder and the DPT decoder, the second phase occurring after the first phase: receive, from a plurality of data sources, polarized image data and ground truth depth; train a second depth model using the received polarized image data and the ground truth data; input the set of initial depth estimates into the second depth model; refine, using the one or more processors, the initial depth estimates based on the trained second depth model and input; and output refined depth estimates.

[0135] F. A system for training a depth estimation model, the system comprising: one or more processors; and a non-transitory machine readable medium containing program instructions that are executable by the one or more processors to perform a method comprising: receiving, using one or more processors, an image data set from a plurality of simulators; training, using the one or more processors, a first depth model using the image data set; generating, using the one or more processors, a set of initial depth estimates using the first depth model; receiving, using the one or more processors, polarized image data and ground truth depth from a plurality of data sources; training, using the one or more processors, a second depth model using the received polarized image data and the ground truth data; inputting, using the one or more processors, the set of initial depth estimates into the second depth model; refining, using the one or more processors, the initial depth estimates based on the training and inputting; and outputting, using the one or more processors, refined depth estimates.

[0136] Each of embodiments D, E, and F may have one or more of the following additional elements in any combination:PCT / US25 / 44748 03 September 2025 (03.09.2025)

[0137] Element 50: wherein the method comprises a first phase using a first vision transformer (ViT) encoder, wherein the first phase includes: receiving the image data set from the plurality of simulators; training the first depth model using the image data set; generating the set of initial depth estimates using the first depth model.

[0138] Element 51 : wherein the method comprises a second phase using a second vision transformer (ViT) encoder, the second phase being after the first phase, the second pass including refining the initial depth estimates based on the training and inputting; and outputting the refined depth estimates.

[0139] Element 52: wherein the method comprises using a Dense Prediction transformer (DPT) decoder.

[0140] Element 53: wherein the first phase further comprises selecting the first depth model, wherein the selection is based on a real-time inference capability of the depth model.

[0141] Element 54: wherein the first phase further comprises receiving and preparing the image data set prior to generation of the set of initial depth estimates.

[0142] Element 55: wherein preparing the image data set includes resizing a plurality of images of the image data set.

[0143] Element 56: wherein the first phase further comprises applying a depthclamping strategy to the initial depth estimates prior to the second phase.

[0144] Element 57: wherein applying the depth-clamping strategy includes capping all pixels in the image data at a pre-specified maximum depth value to penalize incorrect predictions.

[0145] Element 58: wherein the second depth model is a polarization correction model.

[0146] Element 59: wherein the second depth model and the first depth model are the same.

[0147] Element 60: wherein the image data set includes indoor and outdoor image data.

[0148] Element 61 : wherein the plurality of data sources includes polarizationsimulating simulators.PCT / US25 / 44748 03 September 2025 (03.09.2025)

[0149] Element 62: wherein the plurality of data sources includes non-polarization simulating simulators.

[0150] Element 63: wherein the plurality of data sources include a multi-aperture optical camera system configured to aggregate sensor data from multiple optical sensors.

[0151] Element 64: wherein the second phase further comprises deriving polarized image data including Angle of Linear Polarization (AoLP) and Degree of Linear Polarization (DoLP) based on stereoscopic information sensed by the multi-aperture optical camera system.

[0152] Element 65: wherein the second phase further comprises deriving polarized image data including Angle of Linear Polarization (AoLP) and Degree of Linear Polarization (DoLP) by projecting each of the plurality of polarized images of the polarized image data to a 0° polarized image space using the initial depth estimates.

[0153] Element 66: wherein each optical sensor of the multi-aperture optical camera system includes a unique polarization filter.

[0154] Element 67: wherein the plurality of data sources include four optical sensors having apertures of 0°, 45°, 90°, and 135°.

[0155] Element 68: wherein the second phase further comprises polarized image data including Angle of Linear Polarization (AoLP) and Degree of Linear Polarization (DoLP) based on stereoscopic information corresponding to the four optical sensors.

[0156] Element 69: wherein the second phase further comprises deriving polarized image data including Angle of Linear Polarization (AoLP) and Degree of Linear Polarization (DoLP) by projecting each of the plurality of polarized images of the polarized image data to a 0° polarized image space using the initial depth estimates.

[0157] Element 70: wherein each of the four optical sensors includes a unique polarization filter.

[0158] Element 71 : wherein the initial depth estimates are generated using domain-specific data.

[0159] Element 72: wherein the initial depth estimates are generated using pretrained weights.

[0160] Element 73: wherein the image data set comprises 3-channel image data.PCT / US25 / 44748 03 September 2025 (03.09.2025)

[0161] Element 74: wherein the instructions are configured to further cause the one or more processors to: in the first phase, select the first depth model, wherein the selection is based on a real-time inference capability of the depth model.

[0162] Element 75: wherein the instructions are configured to further cause the one or more processors to: in the first phase, receive and prepare the image data set prior to generation of the set of initial depth estimates.

[0163] Element 76: wherein preparing the image data set includes resizing a plurality of images of the image data set.

[0164] Element 77: wherein the instructions are configured to further cause the one or more processors to: in the first phase, apply a depth-clamping strategy to the initial depth estimates prior to refining the initial depth estimates.

[0165] Element 78: wherein applying the depth-clamping strategy includes capping all pixels at a pre-specified maximum depth value to penalize incorrect predictions.

[0166] Element 79: wherein the second depth model is a polarization correction model.

[0167] Element 80: wherein the second depth model and the first depth model are the same.

[0168] Element 81 : wherein the image data set includes indoor and outdoor image data.

[0169] Element 82: wherein the plurality of data sources includes polarizationsimulating simulators.

[0170] Element 83: wherein the plurality of data sources includes non-polarization simulating simulators.

[0171] Element 84: wherein the plurality of data sources include a multi-aperture optical camera system configured to aggregate sensor data from multiple optical sensors.

[0172] Element 85: wherein the instructions are configured to further cause the one or more processors to: in the second phase, derive polarized image data including Angle of Linear Polarization (AoLP) and Degree of Linear Polarization (DoLP) based on stereoscopic information sensed by the multi-aperture optical camera system.

[0173] Element 86: wherein the instructions are configured to further cause the one or more processors to: in the second phase, derive polarized image data including AnglePCT / US25 / 44748 03 September 2025 (03.09.2025) of Linear Polarization (AoLP) and Degree of Linear Polarization (DoLP) by projecting each of the plurality of polarized images of the polarized image data to a 0° polarized image space using the initial depth estimates.

[0174] Element 87: wherein each optical sensor of the multi-aperture optical camera system includes a unique polarization filter applied to each aperture.

[0175] Element 88: wherein the plurality of data sources include four optical sensors having apertures of 0°, 45°, 90°, and 135°.

[0176] Element 89: wherein the instructions are configured to further cause the one or more processors to: in the second phase, derive polarized image data including Angle of Linear Polarization (AoLP) and Degree of Linear Polarization (DoLP) based on stereoscopic information corresponding to the four optical sensors.

[0177] Element 90: wherein the instructions are configured to further cause the one or more processors to: in the second phase, derive polarized image data including Angle of Linear Polarization (AoLP) and Degree of Linear Polarization (DoLP) by projecting each of the plurality of polarized images of the polarized image data to a 0° polarized image space using the initial depth estimates.

[0178] Element 91 : wherein each of the four optical sensors includes a unique polarization filter.

[0179] Element 92: wherein the initial depth estimates are determined using domain-specific data.

[0180] Element 93: wherein the initial depth estimates are determined using pretrained weights.

[0181] Element 94: wherein the image data set comprises 3-channel image data.

[0182] Element 95: wherein the first phase further comprises selecting the first depth model, wherein the selection is based on a real-time inference capability of the depth model.

[0183] Element 96: wherein the first phase further comprises receiving and preparing the image data set prior to generation of the set of initial depth estimates.

[0184] Element 97: wherein preparing the image data set includes resizing a plurality of images of the image data set.PCT / US25 / 44748 03 September 2025 (03.09.2025)

[0185] Element 98: wherein the first phase further comprises applying a depthclamping strategy to the initial depth estimates prior to the second phase.

[0186] Element 99: wherein applying the depth-clamping strategy includes capping all pixels in the image data at a pre-specified maximum depth value to penalize incorrect predictions.

[0187] Element 100: wherein the second depth model is a polarization correction model.

[0188] Element 101 : wherein the second depth model and the first depth model are the same.

[0189] Element 102: wherein the image data set includes indoor and outdoor image data.

[0190] Element 103: wherein the plurality of data sources includes polarizationsimulating simulators.

[0191] Element 104: wherein the plurality of data sources comprises one or more optical instruments.

[0192] Element 105: one or more optical instruments includes cameras, time of flight cameras, structured illumination, light detection and ranging systems (LiDARs), laser range finders, proximity sensors, or multi-aperture optical camera systems.

[0193] Element 106. wherein the one or more optical instruments includes a multiaperture optical camera system having a unique polarization filter applied to each aperture.

[0194] Element 107: wherein the one or more optical instruments includes four optical sensors having apertures of 0°, 45°, 90°, and 135°.

[0195] Element 108: wherein the one or more processors are coupled to at least one interface component configured to combine data output from the one or more optical instruments into individual serial links.

[0196] Element 109: wherein the plurality of data sources includes non-polarization simulating simulators.

[0197] Element 110: wherein the plurality of data sources include a multi-aperture optical camera system configured to aggregate sensor data from multiple optical sensors.PCT / US25 / 44748 03 September 2025 (03.09.2025)

[0198] Element 111 : wherein the second phase further comprises deriving polarized image data including Angle of Linear Polarization (AoLP) and Degree of Linear Polarization (DoLP) based on stereoscopic information sensed by the multi-aperture optical camera system.

[0199] Element 112: wherein the second phase further comprises deriving polarized image data including Angle of Linear Polarization (AoLP) and Degree of Linear Polarization (DoLP) by projecting each of the plurality of polarized images of the polarized image data to a 0° polarized image space using the initial depth estimates.

[0200] Element 113: wherein each optical sensor of the multi-aperture optical camera system includes a unique polarization filter.

[0201] Element 114: wherein the second phase further comprises polarized image data including Angle of Linear Polarization (AoLP) and Degree of Linear Polarization (DoLP) based on stereoscopic information corresponding to the four optical sensors.

[0202] Element 115: wherein the second phase further comprises deriving polarized image data including Angle of Linear Polarization (AoLP) and Degree of Linear Polarization (DoLP) by projecting each of the plurality of polarized images of the polarized image data to a 0° polarized image space using the initial depth estimates.

[0203] Element 116: wherein each of the four optical sensors includes a unique polarization filter.

[0204] Element 117: wherein the plurality of data sources include four optical sensors having apertures of 0°, 45°, 90°, and 135°.

[0205] Element 118: further comprising one or more sensors including ultrasonic sensors, motion sensors, light sensors, infrared sensors, or custom sensors.

[0206] Element 119: wherein the one or more processors includes one or more of a central processing unit, digital signal processor, Application Specific Integrated Circuit, graphics processing unit, or a combination thereof.

[0207] Element 120: wherein the system is an autonomous vehicle, an autonomous mobile robot, an unmanned vehicle, or a bipedal robot.

[0208] Element 121 : wherein the initial depth estimates are generated using domain-specific data.PCT / US25 / 44748 03 September 2025 (03.09.2025)

[0209] Element 122: wherein the initial depth estimates are generated using pretrained weights.

[0210] Element 123: wherein the image data set comprises 3-channel image data.

[0211] G. A system for improved depth estimation, the system comprising: one or more processors; and a non-transitory machine readable medium containing program instructions that are executable by the one or more processors to perform a method comprising: receiving, using one or more processors, polarized images from a multiaperture camera; receiving, using the one or more processors, an initial depth map for the image using a depth model; and generating a metric depth map corresponding to the images using the initial depth map and the polarized images.

[0212] Embodiment G may have one or more of the following additional elements in any combination:

[0213] Element 124: receiving, using one or more processors, an image.

[0214] Element 125: generating and outputting, using the one or more processors, an initial depth map for the image using a depth model.

[0215] Element 126: calculating and outputting, using the one or more processors, polarization information of the image using the polarized images and the initial depth map.

[0216] Each of embodiments A, B, C, D, E, F, and G may have one or more of the following additional elements in any combination:

[0217] Element 127: wherein the system includes autonomous vehicles, autonomous mobile robots, unmanned autonomous vehicles, bipedal robots, aerial drones, industrial robots, autonomous robots, pet robots, robotic vacuums, pet robots, self-driving cars, healthcare related robots

[0218] Element 128: wherein the autonomous vehicles includes one or more of motor vehicles, aerial vehicles, aquatic vehicles, and / or aerospace vehicles.

[0219] Element 129: wherein autonomous robots includes one or more of warehouse robots, sorting robots, factory collaboration robots, retail robots, or hospitality robots.

[0220] Element 130: wherein healthcare related robots includes hospital delivery robots, surgical assistant robots, social companion robots.PCT / US25 / 44748 03 September 2025 (03.09.2025)

[0221] Unless otherwise indicated, all numbers expressing quantities and the like in the present specification and associated claims are to be understood as being modified in all instances by the term “about.” Accordingly, unless indicated to the contrary, the numerical parameters set forth in the specification and attached claims are approximations that may vary depending upon the desired properties sought to be obtained by the embodiments of the present invention. At the very least, and not as an attempt to limit the application of the doctrine of equivalents to the scope of the claim, each numerical parameter should at least be construed in light of the number of reported significant digits and by applying ordinary rounding techniques.

[0222] One or more illustrative embodiments incorporating various features are presented herein. Not all features of a physical implementation are described or shown in this application for the sake of clarity. It is understood that in the development of a physical embodiment incorporating the embodiments of the present invention, numerous implementation-specific decisions must be made to achieve the developer's goals, such as compliance with system-related, business-related, government-related and other constraints, which vary by implementation and from time to time. While a developer's efforts might be time-consuming, such efforts would be, nevertheless, a routine undertaking for those of ordinary skill in the art and having benefit of this disclosure.

[0223] While various systems, tools and methods are described herein in terms of “comprising” various components or steps, the systems, tools and methods can also “consist essentially of” or “consist of” the various components and steps.

[0224] As used herein, the phrase “at least one of” preceding a series of items, with the terms “and” or “or” to separate any of the items, modifies the list as a whole, rather than each member of the list (i.e. , each item). The phrase “at least one of” allows a meaning that includes at least one of any one of the items, and / or at least one of any combination of the items, and / or at least one of each of the items. By way of example, the phrases “at least one of A, B, and C” or “at least one of A, B, or C” each refer to only A, only B, or only C; any combination of A, B, and C; and / or at least one of each of A, B, and C.

[0225] Therefore, the disclosed systems, tools and methods are well adapted to attain the ends and advantages mentioned as well as those that are inherent therein. ThePCT / US25 / 44748 03 September 2025 (03.09.2025) particular embodiments disclosed above are illustrative only, as the teachings of the present disclosure may be modified and practiced in different but equivalent manners apparent to those skilled in the art having the benefit of the teachings herein. Furthermore, no limitations are intended to the details of construction or design herein shown, other than as described in the claims below. It is therefore evident that the particular illustrative embodiments disclosed above may be altered, combined, or modified and all such variations are considered within the scope of the present disclosure. The systems, tools and methods illustratively disclosed herein may suitably be practiced in the absence of any element that is not specifically disclosed herein and / or any optional element disclosed herein. While systems, tools and methods are described in terms of “comprising,” “containing,” or “including” various components or steps, the systems, tools and methods can also “consist essentially of” or “consist of” the various components and steps. All numbers and ranges disclosed above may vary by some amount. Whenever a numerical range with a lower limit and an upper limit is disclosed, any number and any included range falling within the range is specifically disclosed. In particular, every range of values (of the form, “from about a to about b,” or, equivalently, “from approximately a to b,” or, equivalently, “from approximately a-b”) disclosed herein is to be understood to set forth every number and range encompassed within the broader range of values. Also, the terms in the claims have their plain, ordinary meaning unless otherwise explicitly and clearly defined by the patentee. Moreover, the indefinite articles “a” or “an,” as used in the claims, are defined herein to mean one or more than one of the elements that it introduces. If there is any conflict in the usages of a word or term in this specification and one or more patent or other documents that may be incorporated herein by reference, the definitions that are consistent with this specification should be adopted.

[0226] While the above description contains many specific embodiments of the invention, these should not be construed as limitations on the scope of the invention, but rather as an example of one embodiment thereof. Accordingly, the scope of the inventions described herein should be determined based on the specific embodiments illustrated.

Claims

CLAIMSWhat is claimed is:1 . A system for improved depth estimation, the system comprising: one or more processors; and a non-transitory machine readable medium containing program instructions that are executable by the one or more processors to perform a method comprising: receiving, using one or more processors, an image; generating and outputting, using the one or more processors, an initial depth map for the image using a depth model; receiving, using the one or more processors, polarized images corresponding to the image; calculating and outputting, using the one or more processors, polarization information of the image using the polarized images and the initial depth map; and fusing the initial depth map and the polarization information to generate a metric depth map.

2. The system of claim 1 , wherein the method comprises a first pass using a first vision transformer (ViT) encoder, wherein the first pass includes receiving the images and generating and outputting the initial depth map.

3. The system of claim 2, wherein the method comprises a second pass using a second vision transformer (ViT) encoder, the second pass being after the first pass, the second pass including: receiving the polarized images corresponding to the image; calculating and outputting the polarization information of the image; and fusing the initial depth map and the polarization information to generate the metric depth map.

4. The system of claim 3, wherein the first pass and the second pass are performed using a vision transformer (ViT) encoder.

5. The system of claim 1 , wherein the initial depth map and the polarization information is fused within a Dense Prediction transformer (DPT).

6. The system of claim 1 , wherein the polarization information comprises an Angle of Linear Polarization (AoLP) and a Degree of Linear Polarization (DoLP) of the image.

7. The system of claim 1 , wherein the image is received from a multi-aperture optical camera system configured to aggregate sensor data from multiple optical sensors.

8. The system of claim 7, wherein the instructions are configured to further cause the one or more processors to: calculate the polarization information based on stereoscopic information sensed by the multi-aperture optical camera system.

9. The system of claim 7, wherein the instructions are configured to further cause the one or more processors to: calculate the polarization information by projecting each of the polarized images to a 0° polarized image space using the initial depth map.

10. The system of claim 6, wherein the image is received from a multi-aperture optical camera system configured to aggregate sensor data from multiple optical sensors.11 . The system of claim 10, wherein the instructions are configured to further cause the one or more processors to: calculate the AoLP and the DoLP based on stereoscopic information sensed by the multi-aperture optical camera system.

12. The system of claim 10, wherein the instructions are configured to further cause the one or more processors to: calculate the AoLP and the DoLP by projecting each of the polarized images to a 0° polarized image space using the initial depth map.

13. The system of claim 7, wherein each optical sensor of the multi-aperture optical camera system includes a unique polarization filter applied to each aperture.

14. The system of claim 6, wherein the image is received from four optical sensors having apertures of 0°, 45°, 90°, and 135°.

15. The system of claim 14, wherein the instructions are configured to further cause the one or more processors to: calculate the AoLP and the DoLP based on stereoscopic information corresponding to the four optical sensors.

16. The system of claim 14, wherein the instructions are configured to further cause the one or more processors to: calculate the AoLP and the DoLP by projecting each of the polarized images to a 0° polarized image space using the initial depth map.

17. The system of claim 14, wherein each of the four optical sensors includes a unique polarization filter.

18. An improved method for generating a metric depth map of an image, the method comprising: receiving, using one or more processors, an image; generating and outputting, using the one or more processors, an initial depth map for the image using a depth model; receiving, using the one or more processors, polarized images corresponding to the image; calculating and outputting, using the one or more processors, polarization information of the image using the polarized images and the initial depth map; and fusing the initial depth map and the polarization information to generate a metric depth map.

19. The method of claim 18, comprising a first pass using a first vision transformer (ViT) encoder, wherein the first pass includes receiving the images and generating and outputting the initial depth map.

20. The method of claim 19, comprising a second pass using a second vision transformer (ViT) encoder, the second pass being after the first pass, the second pass including: receiving the polarized images corresponding to the image; calculating and outputting the polarization information of the image; and fusing the initial depth map and the polarization information to generate themetric depth map.

21. The method of claim 20, wherein the first pass and the second pass are performed using a vision transformer (ViT) encoder.

22. The method of claim 1 , wherein the initial depth map and the polarization information is fused within a Dense Prediction transformer (DPT).

23. The method of claim 1 , wherein the polarization information comprises an Angle of Linear Polarization (AoLP) and a Degree of Linear Polarization (DoLP) of the image.

24. The method of claim 18, wherein the image is received from a multi-aperture optical camera system configured to aggregate sensor data from multiple optical sensors.

25. The method of claim 24, further comprising calculating the polarization information based on stereoscopic information sensed by the multi-aperture optical camera system.

26. The method of claim 24, further comprising calculating the polarization information by projecting each of the polarized images to a 0° polarized image space using the initial depth map.

27. The method of claim 23, wherein the image is received from a multi-aperture optical camera system configured to aggregate sensor data from multiple optical sensors.

28. The method of claim 27, wherein the Angle of Linear Polarization (AoLP) and Degree of Linear Polarization (DoLP) is calculated based on stereoscopic information sensed by the multi-aperture optical camera system.

29. The method of claim 27, wherein the AoLP and the DoLP is calculated by projecting each of the polarized images to a 0° polarized image space using the initial depth map.

30. The method of claim 27, wherein each optical sensor of the multi-aperture optical camera system includes a unique polarization filter.31 . The method of claim 30, wherein the image is received from optical sensors having apertures of 0°, 45°, 90°, and 135°.

32. The method of claim 31 , wherein the AoLP and the DoLP is calculated based on stereoscopic information corresponding to the four optical sensors.

33. The method of claim 31 , wherein the AoLP and the DoLP is calculated by projecting each of the polarized images to a 0° polarized image space using the initial depth map.

34. The method of claim 31 , wherein each of the four optical sensors includes a unique polarization filter.

35. One or more computer-readable non-transitory storage media including instructions that, when executed by one or more processors, are configured to cause the one or more processors to: receive, using one or more processors, an image; generate and output, using the one or more processors, an initial depth map for the image using a depth model; receive, using the one or more processors, polarized images corresponding to the image; calculate and output, using the one or more processors, polarization information of the image using the polarized images and the initial depth map; and fuse the initial depth map and the polarization information to generate a metric depth map.

36. The one or more computer-readable non-transitory storage media of claim 35, further comprising a first pass using a first vision transformer (ViT) encoder, wherein the first pass includes receiving the images and generating and outputting the initial depth map.

37. The one or more computer-readable non-transitory storage media of claim36, further comprising a second pass using a second vision transformer (ViT) encoder, the second pass being after the first pass, the second pass including: receiving the polarized images corresponding to the image; calculating and outputting the polarization information of the image; and fusing the initial depth map and the polarization information to generate the metric depth map.

38. The one or more computer-readable non-transitory storage media of claim37, wherein the first pass and the second pass are performed using a vision transformer (ViT) encoder.

39. The one or more computer-readable non-transitory storage media of claim 35, wherein the initial depth map and the polarization information is fused within a Dense Prediction transformer (DPT).

40. The one or more computer-readable non-transitory storage media of claim 35, wherein the polarization information comprises an Angle of Linear Polarization (AoLP) and a Degree of Linear Polarization (DoLP) of the image.41 . The one or more computer-readable non-transitory storage media of claim 35, wherein the image is received from a multi-aperture optical camera system configured to aggregate sensor data from multiple optical sensors.

42. The one or more computer-readable non-transitory storage media of claim 41 , wherein the instructions are configured to further cause the one or more processors to: calculate the polarization information based on stereoscopic information sensed by the multi-aperture optical camera system.

43. The one or more computer-readable non-transitory storage media of claim 41 , wherein the instructions are configured to further cause the one or more processors to: calculate the polarization information by projecting each of the polarized images to a 0° polarized image space using the initial depth map.

44. The one or more computer-readable non-transitory storage media of claim 40, wherein the image is received from a multi-aperture optical camera system configured to aggregate sensor data from multiple optical sensors.

45. The one or more computer-readable non-transitory storage media of claim 44, wherein the instructions are configured to further cause the one or more processors to: calculate the AoLP and the DoLP based on stereoscopic information sensed by the multi-aperture optical camera system.

46. The one or more computer-readable non-transitory storage media of claim 44, wherein the instructions are configured to further cause the one or more processors to: calculate the AoLP and the DoLP by projecting each of the polarized images to a 0° polarized image space using the initial depth map.

47. The one or more computer-readable non-transitory storage media of claim 44, wherein each optical sensor of the multi-aperture optical camera system includes a unique polarization filter applied to each aperture.

48. The one or more computer-readable non-transitory storage media of claim 40, wherein the image is received from four optical sensors having apertures of 0°, 45°, 90°, and 135°.

49. The one or more computer-readable non-transitory storage media of claim 48, wherein the instructions are configured to further cause the one or more processors to: calculate the AoLP and the DoLP based on stereoscopic information corresponding to the four optical sensors.

50. The one or more computer-readable non-transitory storage media of claim 48, wherein the instructions are configured to further cause the one or more processors to: calculate the AoLP and the DoLP by projecting each of the polarized images to a 0° polarized image space using the initial depth map.51 . The one or more computer-readable non-transitory storage media of claim 48, wherein each of the four optical sensors includes a unique polarization filter.