System and method for an improved camera system using filters and machine learning to estimate depth

The camera system with directional optics and ML models addresses the inefficiencies of pseudo-LIDAR systems by optimizing light angles and filters, achieving efficient and accurate depth estimation with reduced complexity and hardware.

JP7855905B2Active Publication Date: 2026-05-11TOYOTA JIDOSHA KK
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
TOYOTA JIDOSHA KK
Filing Date
2022-04-13
Publication Date
2026-05-11

AI Technical Summary

Technical Problem

Pseudo-LIDAR systems using multiple cameras and sensors for depth estimation are computationally intensive and complex, leading to inefficiencies in accurately estimating depth, with increased hardware size and latency.

Method used

A camera system utilizing directional optics and machine learning models processes image data through lenses and detectors with varied light angles, employing metasurfaces and filters to reduce computational complexity and improve depth estimation accuracy.

Benefits of technology

The system achieves efficient and accurate depth estimation by reducing computational load and hardware size, producing improved image data for subsequent ML tasks, similar to LIDAR systems but with simpler hardware and processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007855905000001
    Figure 0007855905000001
  • Figure 0007855905000002
    Figure 0007855905000002
  • Figure 0007855905000003
    Figure 0007855905000003
Patent Text Reader

Abstract

To provide systems and methods for an improved camera system using filters and machine learning to estimate depth.SOLUTION: System, methods, and other embodiments described herein relate to estimating depth using a machine learning (ML) model. In one embodiment, a method includes acquiring image data according to criteria from a detector that uses a lens to resolve multiple angles of light per section of the detector. The method also includes mapping a kernel to the image data according to a view associated with the section and a size of the kernel. The method also includes processing the image data using the ML model to produce the depth according to the size of the kernel.SELECTED DRAWING: Figure 6
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The subject matter described herein generally relates to camera systems, and more specifically, to improved camera systems including directional optics, and machine learning models for estimating depth.

Background Art

[0002] Vehicles may be equipped with sensors that facilitate perceiving additional aspects of other vehicles, obstacles, pedestrians, and the surrounding environment. For example, a vehicle may be equipped with a LIDAR (light detection and ranging) sensor that uses light to scan the surrounding environment, and the logic circuitry associated with the LIDAR analyzes the acquired data to detect the location of objects and other features of the scene. In a further example, additional / alternative sensors, such as a camera system, may be implemented to obtain information about the surrounding environment, from which the system derives an awareness of the situation of the surrounding environment. This sensor data can be useful in various environments to improve the perception of the surrounding environment so that a system, such as an autonomous driving system, can understand the recorded situation, make accurate plans, and navigate accordingly.

[0003] Generally, the more a vehicle develops an awareness of its surroundings, the more it can complement the driver with information to assist driving, and / or the autonomous system can better control the vehicle to avoid danger. Systems that use LIDAR to detect objects are best suited for long distances. Therefore, vehicles may use pseudo-LIDAR systems to detect objects using images processed by systems that use multiple cameras and sensors for both short and long distances. However, pseudo-LIDAR systems that rely on multiple cameras and sensors can introduce computational complexity. Also, pseudo-LIDAR systems can use images that change in time and space to create a spatial distribution or point cloud related to the estimated depth, similar to LIDAR systems. Systems that process image data for an accurate spatial distribution can sometimes be complex.

[0004] Furthermore, pseudo-LIDAR systems may acquire images from multiple cameras to estimate depth. They may also modify images from multiple cameras using machine processing to find image overlaps. For example, image overlaps might be stereographs with two or more images sharing corresponding image points. However, pseudo-LIDAR systems that search for image overlaps are time-consuming and computationally intensive. [Overview of the Initiative]

[0005] In one embodiment, the example of a system and method relates to an improved pseudo-light detection and ranging method using an improved camera system including directional optical components and an ML (machine learning) model for depth estimation. In various implementations, pseudo-LIDAR systems are computationally intensive when accurately detecting objects in a scene, as they combine data from multiple sensors or cameras to create a spatial point distribution. Furthermore, the hardware of pseudo-LIDAR using multiple sensors can increase the size of the components, processing tasks, and latency for depth estimation. Thus, pseudo-LIDAR systems can encounter difficulties in efficiently and accurately estimating depth, leading to frustration. Therefore, in one embodiment, the camera system reduces the computation for estimating the depth of the scene by using an ML model, hardware, and input from a limiting sensor to vary the angle of the light wave relative to the image. The output of the camera system may be wide field-of-view image data due to the combination of redundant information in the scene. The system may vary the angle of the light wave according to lens parameters optimized for depth estimation. In addition, the system may use ML models to reduce computation by processing a portion of the image data generated by the lens and detector. The system may also use processed image data from the ML models to classify and estimate the depth of objects in the scene near the camera system.

[0006] Furthermore, the camera system may vary a specific angle of the light wave to subsequently estimate depth using inverted or graded lenses and independent filtering for each pixel of the detector. To improve object detection, the camera system may redirect the light wave associated with the reduced-resolution image for each pixel of the detector array. The output of the camera system may be improved image data containing the object, simplifying subsequent ML or tasks for depth estimation.

[0007] In addition, camera systems may use inverting or shallow-gradient lenses to filter light related to an object by dividing the detector into regions, and then vary a specific angle of the light wave to estimate the depth. A region may be related to one or more pixels. For example, a camera system may use quadrant filtering with a detector divided by quadrants representing different focal areas, and vary the angle of the light wave related to the image.

[0008] Vehicles may be equipped with camera systems that use pixel-by-pixel or quadrant-by-quadrant filtering, depending on efficiency or quality requirements. Furthermore, camera systems may filter light by wavelength for color images before further filtering. In one approach, camera systems may use resonant waveguide gratings (RWGs) on the light to vary its angles. Camera systems may use RWGs as bandpass filters to transmit the changed angles of light of that wavelength to the detector pixels or quadrants. RWGs can improve image detection by outputting image data of a color image containing the object, simplifying subsequent ML or tasks for depth estimation.

[0009] In one embodiment, a camera system for estimating depth using an ML model is disclosed. The camera system includes memory communicably coupled to a processor. The memory stores an acquisition module, which, when executed by the processor, includes instructions that cause the processor to acquire image data according to a criterion from a detector that uses lenses to vary multiple angles of light for each region of the detector. The memory also stores a decision module, which, when executed by the processor, includes instructions that cause the processor to process the image data using an ML model to map kernels to the image data according to the size of the kernels and the views associated with the regions, and to produce depth according to the size of the kernels.

[0010] In another embodiment, a non-temporary computer-readable medium is disclosed that uses an ML model to estimate depth and includes instructions that cause a processor to perform one or more functions when executed by the processor. The instructions include instructions for acquiring image data according to a criterion from a detector that uses lenses to vary multiple angles of light for each region of the detector. The instructions also include instructions for mapping kernels to image data according to the size of the kernels and the views associated with the regions. The instructions also include instructions for processing the image data using an ML model to produce depth according to the size of the kernels.

[0011] In another embodiment, a method for estimating depth using an ML model is disclosed. In one embodiment, the method includes acquiring image data according to a criterion from a detector that uses lenses to vary multiple images of light for each region of the detector. The method also includes mapping kernels to the image data according to the view and kernel size associated with the region. The method also includes processing the image data using an ML model to produce depth according to the kernel size. [Brief explanation of the drawing]

[0012] The attached drawings are incorporated into and constitute part of the patent specification and illustrate various systems, methods, and other embodiments of the present disclosure. It will be understood that the boundaries of elements shown in the drawings (e.g., boxes, groups of boxes, or other shapes) represent one embodiment of the boundary. In some embodiments, one element may be designed as multiple elements, or multiple elements may be designed as a single element. In some embodiments, an element shown as an internal component of another element may be implemented as an external component, and vice versa. Furthermore, elements may not be drawn in exact proportions.

[0013] [Figure 1A] This paper describes a camera system that uses filters to estimate the depth of objects in a scene, and various embodiments of the output of that camera system. [Figure 1B] This paper describes a camera system that uses filters to estimate the depth of objects in a scene, and various embodiments of the output of that camera system. [Figure 1C] This paper describes a camera system that uses filters to estimate the depth of objects in a scene, and various embodiments of the output of that camera system. [Figure 2] This shows a machine learning (ML) model for estimating the depth associated with objects in a scene. [Figure 3] This document presents one embodiment of an estimation system related to improved depth estimation for objects in a scene. [Figure 4] This shows one embodiment of a filter used by a camera system to change the angle of a light wave. [Figure 5A] This describes an embodiment of a camera system that filters areas to estimate the depth associated with objects in a scene. [Figure 5B] This describes an embodiment of a camera system that filters areas to estimate the depth associated with objects in a scene. [Figure 6]One embodiment of a method is presented that includes an estimation system related to determining the depth of objects in a scene. [Figure 7] This shows a camera system that filters light. [Figure 8] One embodiment of a vehicle in which the systems and methods disclosed herein may be implemented is shown. [Modes for carrying out the invention]

[0014] Systems, methods, and other embodiments relating to improvements to camera systems using directional optics, physical filters, and ML models for depth estimation are disclosed herein. The camera system detects objects and reduces costs by optically changing the angle of light waves from multiple views captured in grayscale or color using inverting or shallow-gradient lenses and metasurfaces. The system may change the angle of light waves according to parameters of properties related to the lens or metasurface optimized for depth estimation by adapting to the planarization effect. The system may also optically remove unwanted parallelism according to parameters, thereby improving the capture of light before further processing. Furthermore, the system may use ML models such as deep learning to process image data represented in a smaller two-dimensional matrix than the complete image in order to reduce processing and increase speed.

[0015] ML models can process two-dimensional matrices according to region-by-region (e.g., pixel or area) captures. The system can map kernels using sizes that are optimal for the image data of a given view or angle of a scene. In one approach, the system can process the image data from the ML model to classify, segment, or estimate the depth of objects in the scene near the camera system. For example, the system can use the results of the ML model to then triangulate objects in the scene related to the image data and generate a point cloud using the estimated depth. The point cloud may be similar to the representation produced by a LIDAR system to estimate depth using simpler hardware and an ML model that reduces processing and improves accuracy.

[0016] In color-based capture, camera systems may process, invert, and filter captured light by color wavelength for area-by-area detection. In one approach, camera systems using pixel-by-pixel filtering may have independent filters on pixels to change the angle of the light wave. Image detection may be improved by filtering and combining views based on individual pixels rather than the entire image. Furthermore, color processing may be performed by RWGs, which transmit the light wave to the detector when the light wave matches the wavelength and angle of the respective color filter and metasurface, for precise color tuning.

[0017] Similar to pixel-by-pixel filtering, a camera system can process an image to detect objects using detectors divided into areas, and then vary a specific angle of light waves for subsequent estimation of depth related to the objects. For example, a camera system might use a quadrant filter of size X×X, depending on the number of quadrants arranged for the detectors. In one approach, a camera system might use apertures of different sizes for the quadrants to allow the capture of varying levels of light per quadrant rather than individual pixels. In this way, the camera system can reduce the processing task by capturing the image according to groups of pixels rather than individual pixels.

[0018] Furthermore, the camera system may filter using metamaterials such as metalenses or metasurfaces, manufactured using any one of the following methods: electron beam lithography, roller printing, or photolithography. Metasurfaces provide a substantially flat profile for higher density use within an array of lenses, thereby improving image processing. The transmission profile of a metasurface may contain a desired region of light for transmission to pixels or areas within an angular range. In this way, the camera system uses metasurfaces and area-by-area filtering to detect images from multiple views with improved accuracy and reduced complexity in order to estimate depth.

[0019] Figures 1A to 1C show a camera system that uses a filter to estimate the depth associated with an object in a scene, and various embodiments of the output of that camera system. The camera system 110 or 150 may be incorporated into a vehicle to detect dangerous objects or obstacles within the field-of-view. However, in the various implementations shown herein, the camera system 110 or 150 may be used in any one of a vehicle, a security system, a traffic system, a municipal monitoring system, a portable device, SLAM (simultaneous localization and mapping robotics), camera tracking, SfM (structure from motion), projective geometry, multi-view stereo for volumetric methods, etc. for multi-perspective imaging using a single camera system. In Figure 1A, the camera system 110 can receive the inverted light 120 associated with the scene using the lens 115 and direct it towards the meta-surface 125. The meta-surface 125 may be a lens or a lens system. In one approach, the lens system may include two or more optical elements with one or more apertures or foci. The aperture may include a diaphragm or a pupil. In this way, the camera system 110 using multiple lenses may process different forms of light.

[0020] The metasurface 125 can invert and cancel the parallel effects of the k-vector of the inverted light 120 by filtering or other means for further filtering. In one approach, the metasurface lens may be configured within the camera system 110 to be aligned near the detector array 140, thereby reducing the size and distortion of the system. The detector array 140 may consist of multiple pixels or groups of pixels. In one approach, the metasurface 125 may consist of a photonic bandgap crystal. However, the system may use any lens composition to invert the inverted light 120 into a suitable form for further processing. Furthermore, the lens 115, metasurface 125, metasurface 135, and detector array 140 may be operably connected. As used herein, the term “operably connected” includes direct or indirect connections, and may include connections without direct physical contact.

[0021] In addition, the metasurface 135 may receive the transmitted light 130 from the metasurface 125. In one approach, the metasurface 135 may provide pixel or area-by-area filtering to vary the angle of the transmitted light 130 for image detection. Depending on efficiency or quality requirements, a vehicle may be equipped with a camera system for pixel or quadrant-by-quadrant filtering. Although quadrants may be used in the examples herein, the camera system may use an area of any size to vary the angle of the light wave by filtering. Further, since the light wave takes the form of a plane wave for capture by the detector array 140, the camera system can vary the angle of the light wave. The metasurface 135 can cancel the flattening of the flat wavefront, thereby improving image detection for depth estimation. The close proximity of the metasurface 135 to another filter for varying the angle may result in minimal distortion of the image due to the substantially flat profile of the metasurface. The light wave passing through the filter provides an image for improved detection or capture. For example, the metasurface 135 may transmit light waves or photons to a single pixel by the arrangement of pixel-by-pixel filters 145 at 15° to 30° from the z-axis, varying the angle of the transmitted light 130. However, in the examples shown herein, the camera system may also transmit light waves or photons at 1° to 45° from the normal. The camera system 110 uses the arrangement of independent filters on the pixels to vary the angle of the light wave individually. In this way, the camera system varies the angle of the light wave at the pixel level to improve quality.

[0022] The size of the independent filter may match the size of the pixels of the detector array 140. The camera system 110 improves image capture for depth estimation using pixel-by-pixel filtering by varying the angle of the light individually from multiple viewpoints rather than by the detector area. In this way, a single pixel of the detector array 140 has a varying angle of the light wave for output to an image processor 175 that generates an image from multiple views.

[0023] Similar to the operation of the metasurface 135 described above, the camera system 150 may vary the angle of the captured light waves to detect an image using a metasurface 155 with a gentler gradient to provide pixel-per-pixel or quadrant-per-quadrant filtering. In one approach, a gentle-gradient lens may be used with the metasurface 135, which can function as an imaging lens to capture and modify an image with its axis 15° off-axis with respect to the normal plane of the detector. For pixel-per-pixel capture, the gradient lens of the metasurface 135 may be divided and further divided across the entire surface of the detector array 160. However, the camera system 150 can use quadrant-per-quadrant filtering by dividing the detector array 160 into areas representing groups of pixels. For example, nine interrelated areas could be used, each representing nine different focal areas of the image.

[0024] Because the light waves take the form of plane waves for capture by the detector array 160, it may be necessary to change the angle of the light waves. The metasurface 155 can reverse the "flattening" of the flat wavefront, thereby improving image detection for depth estimation. The close proximity of the metasurface 155 to another filter for changing the angle may result in minimal image distortion due to the substantially flat profile of the metasurface. The light waves passing through the filters provide an image for improved detection or capture.

[0025] In addition, the metasurface 135 or 155 can filter light by using area-by-area filtering of two or more pixels. Areas of two or more pixels may represent different focal areas, views, or image offsets. In one approach, the camera system can use quadrant filters of dimensions X×X, where X is the size of a pixel for a single filter. Filters 170 on the detector or pixel area 165 may include several quadrants, divisions, or regions. In one approach, apertures of different sizes may be used for quadrants, rather than individual pixels, to allow the capture of different levels of light. The camera system can capture multiple views of a scene on a single detector or array of pixels by detecting light waves directed, emitted, or scattered from a specific direction to a specific area, rather than on a pixel-by-pixel basis, to detect the image. However, the system can use pixel-by-pixel processing to produce depth estimates of different qualities.

[0026] Figure 1B shows a metasurface 172 that functions as a lens for a predetermined wavelength of light by manipulating its phase. The unit cells 174 of the metasurface 172 are tuned for pixel or quadrant filtering, changing the angle of light. The unit cells 174 may contain one or more nanostructure elements, represented as rectangles, for directing the light wave and controlling its optical phase relative to the light wave. For example, the unit cells may be tuned to operate according to a corresponding grayscale filter or RGB (red, green, and blue) filter, so that the size or shape of the metasurface adds or removes the appropriate phase required for it to operate at a predetermined wavelength for filtering. Furthermore, unlike inversion lenses, the metasurface is substantially flat. For example, the metasurface may be less than 1 micron in height. Its size as a metasurface material allows for higher density use on detectors.

[0027] Figure 1C shows a system that captures nine images 180 offset from multiple views of a scene containing an object. Although nine images are shown, the system can capture any number of images according to the views present for the scene to estimate depth. Furthermore, as described below, a camera system using directional optics, physical filters, and an ML model can process the images 180 to produce a spatial point distribution or point cloud representation 182 with an improved depth estimate from the improved combined view of the images 180. In one approach, an object in the spatial point distribution or point cloud representation 182 may have a reduced resolution due to the use of nine images from multiple views for a more optimal depth resulting from the increased range of data. Furthermore, the spatial point distribution or point cloud representation 182 may be similar to the representation produced by a LIDAR system to estimate depth with less complexity using camera systems 110 or 150.

[0028] Figure 2 shows an ML model 200 for estimating the depth associated with objects in a scene. The image data may be a two-dimensional representation of images of multiple views of the scene, captured by camera systems 110 or 150. In the case of pixel-by-pixel capture, the two-dimensional matrix may be (rows of pixels) × (columns of pixels), representing the total intensity, RGB values ​​if applicable, and angular information of multiple captured images by pixels. The two-dimensional matrix may exclude information for the z component to reduce processing. In the case of area-by-area processing, the two-dimensional matrix may represent a portion of the multiple captured images (e.g., nine images). In one approach, the image data 210 may also be formatted to represent several images, image height, image width, image depth, etc., to improve the spatial point distribution or point cloud generation.

[0029] Image data 210 is processed by encoder 220 for deep learning, such as by a CNN (convolutional neural network). Convolution is a specific linear operation used in neural networks instead of general matrix multiplication to reduce processing. Thus, at least one layer of encoder 220 may use a CNN to reduce the number of free parameters, allowing the network to go deeper with fewer parameters and in a simpler way. Instead of receiving input from all elements of the previous layer, the convolutional layer receives input from the previous layer or a limited sub-area of ​​the receptive field. In the convolutional layer, the receptive area is smaller than the entire previous layer.

[0030] In the case of a CNN that processes image data, the sub-areas of the original input data within the receptive field grow deeper in the network structure as the convolution operation is reapplied, taking into account the values ​​of specific pixels and surrounding pixels. In one approach, the image data 210 is abstracted after passing through the convolutional layers into a feature map having (number of images) × (height of feature map) × (width of feature map) × (channels of feature map). Furthermore, the convolutional layers of a neural network may be associated with several input and output channels, or hyperparameters. The depth of the convolutional filter or input channel is equal to the number of channels in the input feature map.

[0031] Furthermore, camera systems using more than a single pixel can output color image data using three matrices with similar contrast characteristics to RGB. The color image is represented by three matrices containing values ​​ranging from 1 to 255. The matrices can be composed of (rows of pixels) × (columns of pixels) × (color), or m × n × z. To improve depth estimation, the camera system may also filter using multiple different ray angles represented by (rows of pixels) × (columns of pixels) × (orientation) for processing by the ML model 200, instead of color.

[0032] In one approach, the encoder 220 includes convolutional layers, pooling layers, a ReLU (rectified linear unit), and / or other functional blocks that process the image data 210 separately according to prior training. Once generated within the encoder 220, the low-resolution representation of the feature map is sent to the decoder 230. In one embodiment, the decoder 230 is an extension of the neural network including the encoder 220. In another embodiment, the decoder 230 may be a generative neural network that receives low-resolution representations of the image data 210 from multiple views of the scene to estimate depth. Furthermore, the ML model 200 can use skip connections 240 between the activation blocks of the encoder 220 and the activation blocks of the decoder 230 to facilitate the variation of higher-resolution details.

[0033] The system can use ML model 200 to estimate depth using kernels of a manageable size for CNNs. In one approach, the kernel size may be proportional to the number of views in per-pixel filtering. The kernel can be a linear kernel, a Gaussian kernel, a polynomial kernel, etc. The ML model can process the data according to the kernel method. The kernel may be a user-defined similarity function over pairs of data points in the raw representation. Furthermore, the kernel method can operate in a high-dimensional implicit feature space without computing the data coordinates in that space. In this way, ML models using the kernel method can distinguish features of image data with less computational cost.

[0034] In one approach, a system can run a CNN on an image by taking gradients across the scene image from multiple views. For example, a 9x9 convolution matrix may be used so that the convolution kernel has a manageable size. In this way, the system can use a CNN to efficiently determine the distance to an object without processing the entire image. The system can also collect data from the corners of a two-dimensional image matrix to reduce the offset, thereby improving the depth estimation task. As an example, a pixel-by-pixel system can calculate gradients across columns of pixels x and detect changes in brightness I. As a result, the derivative dI / dx may be used as an indicator of an object. Furthermore, the system may identify viewpoints of multiple views within the kernel, by binning, and across the kernel to provide entirely distinct outputs compared to outputs without that pixel. Information across the scene may be collected by a system in a local region through multiple views using known parameters such as angle and depth. Therefore, a camera system may use a CNN to estimate depth using multiple views of the scene with less complexity and computational cost.

[0035] Figure 3 shows that the estimation system 310 is related to an improved estimation of depth for objects in a scene. The estimation system 310 is shown as including a processor 320 from the vehicle 800 in Figure 8. Thus, the processor 320 may be part of the estimation system 310, the estimation system 310 may include a processor separate from the processor 320 of the vehicle 800, or the estimation system 310 may access the processor 320 through a data bus or another communication path. In one embodiment, the estimation system 310 includes a memory 330 that stores an acquisition module 340 and a determination module 350. The memory 330 is RAM (random-access memory), ROM (read-only memory), a hard disk drive, flash memory, or other suitable memory for storing modules 340 and 350. For example, modules 340 and 350 are computer-readable instructions that cause the processor 320 to perform various functions disclosed herein when executed by the processor 320.

[0036] Furthermore, the acquisition module 340 generally includes instructions that function to control the processor 320 to receive input data from one or more sensors of the vehicle 800. In one embodiment, the input is an observation of one or more objects in the environment closest to the vehicle 800, and / or other conditions about the surroundings. In one embodiment, as specified herein, the acquisition module 340 acquires sensor data 370, which includes at least a camera image.

[0037] Therefore, in one embodiment, the acquisition module 340 controls each sensor to provide data input in the form of sensor data 370. In addition, although the acquisition module 340 has been described as controlling various sensors to provide sensor data 370, in one or more embodiments, the acquisition module 340 may use other techniques, either active or passive, for acquiring the sensor data 370. For example, the acquisition module 340 may passively intercept the sensor data 370 from a stream of electronic information provided by various sensors to additional components inside the vehicle 800. Furthermore, the acquisition module 340 can implement various approaches for fusing data from multiple sensors and / or from sensor data acquired across the entire wireless communication link when providing sensor data 370.

[0038] In one embodiment, the estimation system 310 includes a data store 360. In one embodiment, the data store 360 ​​is a database. In one embodiment, the database is an electronic data structure stored in memory 330 or another data store and consists of routines that can be executed by the processor 320 for purposes such as analyzing the stored data, providing the stored data, and organizing the stored data. Thus, in one embodiment, the data store 360 ​​stores data used by modules 340 and 350 when performing various functions. In one embodiment, the data store 360 ​​includes sensor data 370, along with metadata that characterizes various aspects of the sensor data 370. For example, the metadata may include location coordinates (e.g., longitude and latitude), coordinates or tile identifiers on the map in question, and a timestamp (time / date stamp) from when the separate sensor data 370 was generated.

[0039] In one embodiment, the datastore 360 ​​further includes any one of a criterion 380 and a kernel size 390. In one approach, the kernel size 390 may relate to the number of unit cells or the number of quadrants. Quadrants may be used in the examples herein, but the camera system may use areas of arbitrary size to change the angle of the light wave by filtering. Furthermore, the system may use the criterion 380 to determine whether the image data satisfies predetermined parameters. The image data may represent pixels or areas of multiple images that have been processed so that a network of lenses changes the angle and combined for capture by a single detector. Moreover, the predetermined parameters may relate to resolution, noise level, angular offset, phase or similar. The system may then use an ML model for further processing to estimate depth using the image data according to the kernel size 390 representing a portion of the image data to be processed, e.g., 3x3. For example, in a CNN, the kernel size 390 may represent a portion of the image data to be processed by a layer of the ML model.

[0040] In one embodiment, the acquisition module 340 is further configured to perform additional tasks beyond controlling each sensor to acquire and provide sensor data 370. For example, the acquisition module 340 includes instructions to cause the processor 320 to acquire image data from the camera system 110 or 150 according to a reference 380. The camera system 110 or 150 may use a metasurface or lens to change multiple angles of light for each area and output the image data for further processing.

[0041] In one embodiment, the determination module 350 includes instructions to cause the processor 320 to map kernels to image data according to views or elements associated with areas and kernel size 390. In the case of area-based filtering, before mapping kernels, pixels in one area may be processed together with corresponding pixels in another area to combine different images from a single camera into elements. The determination module 350 may also process the image data to estimate depth according to kernel size 390 and classify objects in the scene related to the image data using an ML model. The classification may be a new observation or label of an object in the scene determined by the ML model from the image data, for example, by feature comparison. For example, the class may be a vehicle, a person, a signal, a light post, a plant, etc.

[0042] Figure 4 shows one embodiment of a filter used by a camera system to change the angle of a light wave. In the case of pixel-by-pixel filtering, the detector may use multiple angle-based filters 410 for filtering. The metasurface or lens 420 may include a maximum limit or highest point 430 where the gradient becomes gentler, and multiple unit cell elements, in order to change the angle of the light wave by reflection and phase change. The metasurface may have a substantially flat profile for use on the detector at higher density and closest to the detector for improved image processing. In one approach, the metasurface may be manufactured using electron beam lithography, roll-to-roll printing, photolithography, etc. In one approach, the multiple angle-based filters 410 may be a gently sloping metasurface lens used by the system to change the angle related to the grayscale image.

[0043] A side view of the metasurface or lens 440 shows the maximum limit to which the gradient becomes gentler. Furthermore, in the case of quadrant filtering, multiple angle-based filters 450 may be placed on-chip, substantially as close to the detector as possible, according to the areas NW (northwest), N (north), NE (northeast), W (west), center, E (east), SW (southwest), S (south), and SE (southeast), in order to capture different parts of multiple images. For example, the area of ​​the angle-based filter 450 may correspond to multiple or groups of pixels for angle-related image capture to reduce complexity. Although quadrants may be used in the examples herein, the camera system may use any area size to change the angle of the light wave by filtering.

[0044] Furthermore, as described below, the system can filter using quadrant-by-quadrant arrangement for gently sloping metasurfaces and lenses equivalent to quadrant-by-quadrant filtering 460. In quadrant filtering, the filter area may be associated with multiple filters to reduce complexity when the angle is changed. Conversely, in the case of pixel-by-pixel arrangement 470, the lens of the gently sloping metasurface may be divided into pixel element units and arranged on the detector. In pixel-by-pixel arrangement 470, the unit cell is adjusted pixel by pixel to change the angle of light. The unit cell may contain one or more nanostructure elements, shown as rectangles, to correct the optical phase with respect to the light wave. In some structures, a camera system using pixel-by-pixel arrangement can generate richer image data and improve depth estimation.

[0045] Figures 5A and 5B show embodiments of a camera system 510 that performs quadrant filtering to estimate depth relative to objects in a scene. For example, in stacking, the quadrant filters 520 may include a gently sloping metasurface or lens through which the angle-resolved light is transmitted for the color filter 530 to process the light waves. In one approach, the color filter 530 may be divided similarly to a standard lens for processing the light waves. A detector 540 can detect the angle-resolved light waves according to the color of the wavelength.

[0046] In the examples shown herein, the color filter may include a red filter with a wavelength of substantially 630 nm (nanometers), a green filter with a wavelength of 530 nm, and a blue filter with a wavelength of 400 nm. However, the color filter can transmit light waves filtered at any wavelength based on color to the detector. In one approach, the filter may be a Bayer filter coupled with other angular bandpass filters.

[0047] Similarly, in another stack, the per-pixel filter 550 may be used by a gently sloping metasurface or lens through which the angle-changing light is transmitted for the color filter 560 to process the light waves. Camera systems using per-pixel arrangements can generate richer image data and improve depth estimation. In one approach, the color filter 560 may be further divided into nine lenses to utilize the same color filter, thereby reducing the number of components. Finally, the detector 570 can detect the angle-changing light waves according to the color of the wavelength.

[0048] In Figure 5B, camera systems 580 and 591 are configured for pixel-level filtering and quadrant-level filtering, respectively, and can change the angle of the light wave. In the case of pixel-level filtering, camera system 580 may utilize an inverting or standard lens 582 to receive light waves associated with an object. The inverting or standard lens 582 may be a single lens or form a lens system. The lens system may include two or more optical elements with one or more apertures. The apertures may include a diaphragm or pupil. A color filter 584 can filter the light waves from the inverting or standard lens 582 by wavelength. For example, camera system 580 may separate the light waves into RGB components for processing.

[0049] The metasurface 586 can filter light waves by removing angle-induced shift, inversion, and / or undo parallel effects at the pixel level, or in relation to the standard lens 582. The unit cells of the metasurface 586 may be tuned to operate within the parameters of the color filter 584 so that the size or shape of the metasurface adds or removes the appropriate phase required for the light tuned through the color filter 584 to operate at a predetermined wavelength. The unit cells of the metasurface 586 may contain one or more nanostructures to correct the optical phase. The metasurface 586 can provide a substantially flat profile corresponding to the size of the color filter 584, allowing for higher density use on the detector and reduced image distortion. For example, the metasurface 586 material may be a photonic bandgap crystal, silicon dioxide, or titanium dioxide. In contrast to other lens systems, the metasurface 586 may be closest to or near the detector 590, resulting in minimal distortion when filtering images. In one approach, the detector 590 may be a detector or pixel array, such as a CMOS (complementary metal-oxide-semiconductor) or CCD (charge-coupled device). An image processor may further process the output of the detector 590 for the subsequent task of estimating depth. The improved image data may represent pixels or areas of multiple images, which have been processed so that the lens network changes angle and combined for capture by a single detector.

[0050] Furthermore, the wavelength-adaptive RWG 588 may transmit light waves to the detector 590 if the light waves match the wavelengths and angles of the color filter 584 and the metasurface 586, respectively. The RWG may operate within the range of the adjusted filter area of ​​the color filter 584. The colored squares in the color filter 584, the metasurface 586, and the RWG 588 may correspond to pixels in the detector 590. In this way, the camera system 580 uses components 584, 586, and 588 to change the desired color and angle combination by combining angular bandpass filtering and color filtering.

[0051] The manufacturing process involves depositing silicon onto the surface of a transparent glass plate to produce RWG (Real-Wooded Image). Since silicon dioxide has a desirable refractive index of 1.45, RWG produced by silicon is sometimes desirable for image capture. In addition, silicon has a refractive index that varies along the wavelength spectrum of 400nm to 700nm, which is suitable for color image capture.

[0052] In one approach, the manufacturing process can use titanium oxide fused with silica to create a RWG for changing the angle of light waves. For example, the RWG may stack two or more filters with a 625 nm red filter to enable transmission at incident angles of 10–25°. The RWG may also suppress waves at other angles for bandpass filtering before detection.

[0053] Furthermore, the RWG may have a minimum transmittance for normal incident light, a high transmittance at an angle of 15°, or a minimum transmittance up to 90° away from the normal. The RWG can limit the light waves transmitted at these incident angles, while still allowing the complete spectrum of the transmitted light waves to be received. Thus, the camera system 580 can vary the desired combination of color and angle by combining angular bandpass filtering and color filtering.

[0054] Furthermore, in the case of quadrant filtering, the camera system 591 may use a divided metasurface or lens 592 that receives light waves at different angles or views related to an object. While the camera system 591 uses quadrants, the detector 590 may be divided into arbitrary groups of pixels to vary the angle related to detecting the depth of objects in the scene of the captured image. A color filter 584 can filter light waves by wavelength. For example, the camera system 591 may separate light waves into RGB components for processing. The metasurface 594 can filter light waves using quadrants by inversion or by removing angle-induced shift, inversion, and / or parallel effects associated with the standard metasurface or lens 592. The metasurface 594 may be a quadrant for a filter of dimensions X × X, where X is the size of the pixels for a single filter, and the quadrant corresponds to the size of the lens of the divided metasurface or lens 592. The metasurface 594 may be several quadrants divided according to the size of the detector 590. In one approach, apertures of different sizes may be used for quadrants, allowing for varying levels of light capture by quadrants rather than pixels.

[0055] Furthermore, the camera system 591 can utilize a wavelength-adaptive RWG 588 that transmits light waves to the detector 590 when the light waves match the wavelength and angle of the color filter 584 and the metasurface 594, respectively. In this way, the camera system uses components 584, 594, and 588 to change the desired color and angle combination by combining angular bandpass filtering and color filtering.

[0056] Additional aspects of the estimation system will be described in relation to Figure 6. In particular, Figure 6 shows one embodiment of a method that includes an estimation system related to determining depth relative to objects in a scene. Method 600 will be described from the perspective of estimation system 310 in Figure 3. While Method 600 will be described in conjunction with estimation system 310, it should be understood that Method 600 is not limited to being implemented within estimation system 310, but rather is an example of a system in which Method 600 can be implemented. In Method 600, the system acquires image data from a camera system. The image data may represent pixels or areas of multiple images that a network of lenses processes to vary angles and combine for capture by a single detector. The metasurface or network of lenses of the camera system varies angles related to the scene represented by the image data. An ML model improves the computational task by processing the image data to estimate depth by analyzing a portion of the scene.

[0057] In the 610, the system acquires image data according to a criterion. The image data may represent multiple images captured or acquired by a single detector, processed so that a metasurface or network of lenses changes angle. The system can use the criterion to determine whether the image data satisfies predetermined parameters. For example, predetermined parameters may relate to resolution, noise level, angular offset, phase, etc.

[0058] In 620, the system maps kernels to views or elements of image data. The system can use an ML model to estimate depth using kernels of a manageable size for the CNN. The kernel method can operate according to the kernel size without computing the coordinates of the image data in a predefined space. In this way, CNNs using the kernel method can distinguish features of image data with less computational cost. In one approach, the kernel size may be proportional to the number of views in pixel-by-pixel filtering. The kernel can be a linear kernel, a Gaussian kernel, a polynomial kernel, etc. For example, the system might use a 9x9 convolution matrix as a manageable size for the convolution kernel. Furthermore, the system can collect data from similar regions of different views of the image matrix to reduce the offset, thereby improving depth estimation. In the case of area-by-area filtering, before mapping the kernel, pixels in one area may be processed together with corresponding pixels in another area to combine different image data from a single camera into elements. Therefore, the system can use a CNN to quickly determine the distance to an object without processing the entire image.

[0059] Furthermore, ML models can process data according to kernel methods. A kernel method may be a user-defined similarity function across pairs of data points in a raw representation. Kernel methods can operate within a high-dimensional implicit feature space without computing data coordinates within that space. In this way, ML models using kernel methods can distinguish features of image data with less computational cost. Additionally, camera systems using pixel-per-pixel capture can output image data relating to multiple views of a scene. Camera systems using area-per-area or quadrant-per-quadrant capture can output image data having multiple elements of a scene. While quadrants may be used in the examples herein, camera systems can use areas of any size to change the angle of light waves through filtering. Thus, the system maps kernels to views or elements of image data.

[0060] In 630, the system processes image data using an ML model, such as deep learning with an encoder and decoder network. To reduce processing, the system can process image data in a two-dimensional matrix smaller than the complete image. The two-dimensional matrix may be (rows of pixels) × (columns of pixels), representing brightness, RGB if applicable, and angular information of the captured image without a z component. In one approach, the ML model can process the two-dimensional matrix according to captures for specific regions (e.g., pixels or areas).

[0061] In one or more embodiments, the system can run a CNN on an image region by region by incorporating gradients across the image. For example, a 9x9 convolution matrix may be used so that the convolution kernel has a manageable size. In this way, the system can use the CNN to efficiently determine the distance to an object without processing the entire image. Furthermore, the system can collect data from the corners of a two-dimensional image matrix to reduce offsets and generate feature maps for further depth estimation tasks. As an example, a pixel-by-pixel system can calculate gradients across columns of pixels x to detect changes in brightness I. As a result, the derivative dI / dx may be used as an indicator of an object. Moreover, the system may identify viewpoints of multiple views within the kernel, by binning, and across the kernel to provide entirely distinct outputs compared to outputs without that pixel. Information across the scene may be collected by the system in local regions through multiple views using known parameters such as angle and depth.

[0062] In 640, the system classifies objects in a scene represented by image data and processed by an ML model. The classification may be a new observation, label, or identification of an object in the scene determined by the ML model from the processed image data. The system can determine the classification by comparing the features of known or unknown objects in the scene. For example, the class may be a vehicle, a person, a traffic light, a plant, etc.

[0063] In version 650, the system determines location information from triangulation of objects in the scene. For example, the system can use an ML model to process image data for instance segmentation of objects in the scene relevant to classification. The system can then determine the location information of the objects according to the results of the instance segmentation and the associated triangulation.

[0064] In 660, the system uses the positional information of triangulated objects in the scene to generate a point cloud using estimated depth. In another example, the system can triangulate multiple objects in the scene according to classification to generate a distribution of spatial points. The point cloud may be similar to the representation generated by a LIDAR system for estimating depth, but using simpler hardware and an ML model that reduces processing while improving accuracy.

[0065] Figure 7 shows a camera system for filtering light. A lens, metalens, or metasurface 710 can filter light waves to focus on the detector 750. Light waves at angle 720 can be processed according to parameters of properties associated with the lens, metalens, or metasurface 710 and output as 730 or 740. For example, the parameters may define the amount of refraction or direction of the light waves at a particular brightness. The system can vary the angle of the light waves according to the parameters of a lens optimized for measuring depth. In the case of RGB, a color filter 760 may be placed as close to the detector 750 as possible to filter the wavelength before changing the angle. In one approach, the filter may be a Bayer filter coupled with other angle bandpass filters.

[0066] Figure 8 shows one embodiment of a vehicle in which the systems and methods disclosed herein may be implemented. As used herein, “vehicle” refers to any form of motor-driven means of transport. In one or more implementations, vehicle 800 is an automobile. While arrangements are described herein in relation to automobiles, it will be understood that embodiments are not limited to automobiles. In some implementations, vehicle 800 may be any robotic device or form of motor-driven means of transport.

[0067] Vehicle 800 also includes various elements. In various embodiments, it will be understood that vehicle 800 may have fewer elements than those shown in Figure 8. Vehicle 800 may have any combination of the various elements shown in Figure 8. Furthermore, vehicle 800 may have additional elements beyond those shown in Figure 8. In some configurations, vehicle 800 may be implemented without one or more of the elements shown in Figure 8. While the various elements are shown to be located inside vehicle 800 in Figure 8, it will be understood that one or more of these elements may be located outside vehicle 800.

[0068] In some examples, the vehicle 800 is configured to selectively switch between different modes of operation / control according to the direction of one or more modules / systems of the vehicle 800. In one approach, the modes include 0, not automated at all; 1, driver assistance; 2, partially automated; 3, conditionally automated; 4, highly automated; and 5, fully automated. In one or more configurations, the vehicle 800 may be configured to operate in a subset of the possible modes.

[0069] In one or more embodiments, the vehicle 800 is an automated vehicle or an autonomous vehicle. As used herein, “automated vehicle” or “autonomous vehicle” refers to a vehicle capable of operating in an autonomous mode (e.g., Category 5, fully automated). “Autonomous mode” refers to using one or more computer systems to navigate and / or operate the vehicle 800 along a travel route and to control the vehicle 800 with minimal or no input from a human driver. In one or more embodiments, the vehicle 800 is highly automated or fully automated. In one embodiment, the vehicle 800 is configured in one or more semi-autonomous operating modes in which one or more computer systems perform part of the navigation and / or operation of the vehicle along a travel route, and the vehicle operator (i.e., driver) provides input to the vehicle to perform part of the navigation and / or operation of the vehicle 800 along the travel route.

[0070] The vehicle 800 may include one or more processors 320. In one or more configurations, the processor 320 may be the main processor of the vehicle 800. For example, the processor 320 may be an ECU (electronic control unit), an ASIC (application-specific integrated circuit), a microprocessor, etc. The vehicle 800 may include one or more data stores 815 for storing one or more types of data. The data stores 815 may include volatile and / or non-volatile memory. Examples of suitable data stores 815 include RAM, flash memory, ROM, PROM (Programmable Read Only Memory), EPROM (Erasable Programmable Read-Only Memory), EEPROM (Electrically Erasable Programmable Read-Only Memory), registers, magnetic disks, optical disks, and hard drives. The data stores 815 may be components of the processor 320, or the data stores 815 may be operablely connected to the processor 320 for use by it. As used throughout this specification, the term “operably connected” may include direct or indirect connections, and may include connections without direct physical contact.

[0071] In one or more configurations, one or more data stores 815 may include map data 816. Map data 816 may include maps of one or more geographic areas. In some examples, map data 816 may include information or data about roads, traffic control equipment, road signs, structures, features and / or landmarks within one or more geographic areas. Map data 816 can take any appropriate form. In some examples, map data 816 may include an aerial view of the area. In some examples, map data 816 may include a ground view of the area, including a 360° ground view. Map data 816 may include measurements, dimensions, distances and / or information about one or more items included in map data 816 and / or other items included in map data 816. Map data 816 may include a digital map with information about the geometry of roads.

[0072] In one or more configurations, map data 816 may include one or more topographic maps 817. A topographic map 817 may include information about the topography, roads, surfaces and / or other features of one or more geographic areas. A topographic map 817 may include elevation data within one or more topographic areas. A topographic map 817 may define one or more ground surfaces, which may include paved roads, unpaved roads, soil and other features defining the ground surface.

[0073] In one or more configurations, map data 816 may include one or more fixed obstacle maps 818. A fixed obstacle map 818 may include information about one or more fixed obstacles located within one or more terrain areas. A “fixed obstacle” is a physical object whose location does not change or substantially change throughout the period, and / or whose size does not change or substantially change throughout the period. Examples of fixed obstacles may include trees, buildings, curbs, fences, guardrails, median strips, utility poles, statues, monuments, signs, benches, furniture, mailboxes, large rocks, or hills. Fixed obstacles may be objects that extend above ground level. One or more fixed obstacles included in a fixed obstacle map 818 may have associated location data, size data, dimension data, material data, and / or other data. A fixed obstacle map 818 may include measurement results, dimensions, distance, and / or information for one or more fixed obstacles. A fixed obstacle map 818 may be of high quality and / or very detailed. The fixed obstacle map 818 can be updated to reflect changes within the area depicted on the map.

[0074] One or more data stores 815 may include sensor data 819. In this context, “sensor data” means any information about sensors equipped on the vehicle 800, including capabilities and other information about such sensors. The vehicle 800 may include a sensor system 820, as shown below. The sensor data 819 may be associated with one or more sensors of the sensor system 820. For example, in one or more configurations, the sensor data 819 may include information about one or more LIDAR sensors 824 of the sensor system 820.

[0075] In some examples, at least a portion of the map data 816 and / or sensor data 819 may be located in one or more data stores 815 mounted on and located in the vehicle 800. Alternatively, or in addition, at least a portion of the map data 816 and / or sensor data 819 may be located in one or more data stores located away from the vehicle 800.

[0076] As described above, the vehicle 800 may include a sensor system 820. The sensor system 820 may include one or more sensors. "Sensor" means a device capable of detecting and / or sensing something. In at least one embodiment, one or more sensors detect and / or sense in real time. As used herein, "real time" means a level of processing responsiveness that is sufficiently immediate for a user or system to sense that a particular process or decision is being made, or that allows a processor to keep pace with some external processing.

[0077] In a configuration in which the sensor system 820 includes multiple sensors, the sensors can function independently, or two or more sensors can function in combination. The sensor system 820 and / or one or more sensors may be operably connected to the processor 320, the data store 815, and / or other elements of the vehicle 800. The sensor system 820 can produce observations about a portion of the environment of the vehicle 800 (e.g., near the vehicle).

[0078] The sensor system 820 may include any suitable type of sensor. Various examples of different types of sensors will be described herein. However, embodiments should be understood not to be limited to the specific sensors described. The sensor system 820 may include one or more vehicle sensors 821. The vehicle sensors 821 may detect information about the vehicle 800 itself. In one or more configurations, the vehicle sensors 821 may be configured to detect changes in the position and orientation of the vehicle 800, for example, based on inertial acceleration. In one or more configurations, the vehicle sensors 821 may include one or more accelerometers, one or more gyroscopes, an IMU (inertial measurement unit), an autonomous navigation system, a GNSS (global navigation satellite system), a GPS (global positioning system), a navigation system 847, and / or other suitable sensors. The vehicle sensors 821 may be configured to detect one or more features of the vehicle 800 and / or how the vehicle 800 is operating. In one or more configurations, the vehicle sensor 821 may include a speedometer for determining the current speed of the vehicle 800.

[0079] Alternatively, or in addition, the sensor system 820 may include one or more environmental sensors 822 configured to acquire data about the environment surrounding the vehicle 800 in which the vehicle 800 is operating. "Environmental data" includes data about the external environment in which the vehicle is located, or one or more parts thereof. For example, one or more environmental sensors 822 may be configured to sense obstacles in at least a part of the external environment of the vehicle 800, and / or data about such obstacles. Such obstacles may be stationary objects and / or moving objects. One or more environmental sensors 822 may be configured to detect other things in the external environment of the vehicle 500, such as lane markers, signs, traffic lights, traffic signs, lane boundaries, crosswalks, curbs closest to the vehicle 500, and off-road objects.

[0080] Various examples of sensors of the sensor system 820 will be described herein. These examples may be part of one or more environmental sensors 822 and / or one or more vehicle sensors 821. However, it will be understood that embodiments are not limited to the specific sensors described.

[0081] As an example, in one or more configurations, the sensor system 820 may include one or more radar sensors 823, LiDAR sensors 824, sonar sensors 825, weather sensors, tactile sensors, position sensors, and / or one or more cameras 826. In one or more configurations, one or more cameras 826 may be HDR (high dynamic range) cameras, stereo or IR (infrared) cameras.

[0082] Vehicle 800 may include an input system 830. The “input system” includes components or configurations, or groups thereof, that enable various entities to input data into the machine. The input system 830 may receive input from the occupants of the vehicle. Vehicle 800 may include an output system 835. The “output system” includes one or more components that facilitate the presentation of data to the occupants of the vehicle.

[0083] Vehicle 800 may include one or more vehicle systems 840. Various examples of one or more vehicle systems 840 are shown in Figure 8. However, vehicle 800 may include more, fewer or different vehicle systems. While certain vehicle systems are defined separately, it should be understood that any or part of a system may be combined or separated in other ways through hardware and / or software within vehicle 800. Vehicle 800 may include a propulsion system 841, a braking system 842, a steering system 843, a throttle system 844, a transmission system 845, a signaling system 846, and / or a navigation system 847. Any of these systems may include one or more devices, components, and / or combinations thereof that are currently known or will be developed later.

[0084] The navigation system 847 may include one or more currently known or later developed devices, applications, and / or combinations thereof configured to determine the geographical location of the vehicle 800 and / or to determine a route for the vehicle 800. The navigation system 847 may include one or more mapping applications for determining a route for the vehicle 800. The navigation system 847 may include a global positioning system, a local positioning system, or a geolocation information system.

[0085] The processor 320 or autonomous driving module 860 may be operablely connected to communicate with various vehicle systems 840 and / or their individual components. For example, the processor 320 and / or autonomous driving module 860 may communicate to send and / or receive information from various vehicle systems 840 to control the movement of the vehicle 800. The processor 320 or autonomous driving module 860 may control some or all of the vehicle systems 840 and therefore may be partially or fully autonomous as defined by SAE (Society of Automotive Engineer) levels 0-5.

[0086] The processor 320 and the autonomous driving module 860 may be operable to control the navigation and operation of the vehicle 800 by controlling one or more vehicle systems 840 and / or their components. For example, when operating in autonomous mode, the processor 320 or the autonomous driving module 860 may control the direction and / or speed of the vehicle 800. The processor 320 or the autonomous driving module 860 may cause the vehicle 800 to accelerate, decelerate and / or change direction. As used herein, “cause” or “causing” means to cause, compel, force, direct, command, guide and / or enable an event or action, or at least put into a state where such an event or action could occur, either directly or indirectly.

[0087] The vehicle 800 may include one or more actuators 850. The actuators 850 may be elements or combinations of elements that can operate to modify one or more vehicle systems 840 or their components in response to receiving signals or other inputs from the processor 320 or the autonomous driving module 860. For example, one or more actuators 850 may include, to name a few, motors, pneumatic actuators, hydraulic pistons, relays, solenoids and / or piezoelectric actuators.

[0088] The vehicle 800 may include one or more modules, at least some of which are described herein. A module may be implemented as computer-readable code that performs one or more of the various operations described herein when executed by the processor 320. One or more modules may be components of the processor 320, or one or more modules may run on and / or be distributed within other processing systems to which the processor 320 is operationally connected. A module may include instructions (e.g., program logic) that are executable by one or more processors 320. Alternatively, or in addition to such instructions, one or more data stores 815 may include such instructions.

[0089] The vehicle 800 may include one or more autonomous driving modules 860. The autonomous driving modules 860 may be configured to receive data from the sensor system 820 and / or from any other type of system capable of capturing information about the vehicle 800 and / or the external environment of the vehicle 800. In one or more configurations, the autonomous driving modules 860 may use such data to generate one or more driving scene models. The autonomous driving modules 860 may determine the position and speed of the vehicle 800. The autonomous driving modules 860 may determine the position of obstacles, obstacles, or other environmental features, including road signs, trees, shrubs, adjacent vehicles, pedestrians, etc.

[0090] The autonomous driving module 860 may be configured to receive and / or determine the location of obstacles in the external environment of the vehicle 800, which the processor 320 and / or one or more modules described herein use to determine the position and orientation of the vehicle 800; the vehicle's position in global coordinates based on signals from multiple satellites; or any other data and / or signals that may be used to determine the current state of the vehicle 800 or the current position of the vehicle 800 in relation to that environment.

[0091] The autonomous driving module 860 may be configured to determine the travel path, the current autonomous driving operation for the vehicle 800, future autonomous driving operations, and / or modifications to the current autonomous driving operation, based on data acquired by the sensor system 820, a driving scene model, and / or data from any other suitable source. “Driving operation” means one or more actions that affect the movement of the vehicle. Examples of driving operations include, to name a few, accelerating, decelerating, braking, turning, moving the vehicle 800 laterally, changing travel lanes, merging into travel lanes, and / or reversing. The autonomous driving module 860 may be configured to perform the determined driving operation. The autonomous driving module 860 may directly or indirectly cause such autonomous driving operations to be performed. As used herein, “cause” or “causing” means to cause, command, guide, and / or enable an event or action, or at least put into a state where such an event or action could occur. The autonomous driving module 860 may be configured to perform various vehicle functions and / or to transmit data to the vehicle 800 or one or more of its systems (e.g., one or more vehicle systems 840), to receive data from the vehicle 800 or one or more of its systems, to interact with the vehicle 800 or one or more of its systems, and / or to control the vehicle 800 or one or more of its systems.

[0092] Detailed embodiments are described herein. However, it should be understood that the disclosed embodiments are intended as examples. Therefore, the specific structural and functional details disclosed herein should not be construed as limitations, but merely as a typical basis for the claims and for teaching those skilled in the art how to apply the embodiments herein in various ways in substantially any appropriately detailed structures. Furthermore, the terms and phrases used herein are not intended to be limiting, but rather to provide an understandable description of possible implementations. Various embodiments are shown in Figures 1-8, but the embodiments are not limited to the structures or applications shown.

[0093] The flowcharts and block diagrams in the drawings illustrate the structure, function, and operation of possible implementations of systems, methods, and computer program products according to various embodiments. In this regard, a block or block diagram in a flowchart may represent a module, segment, or part of code that contains one or more executable instructions for implementing a particular logical function. It should also be noted that in some alternative implementations, the functions mentioned within a block may occur in a different order than the order in which they are mentioned in the drawing. For example, two blocks shown consecutively may actually be executed substantially simultaneously, or they may sometimes be executed in the opposite order depending on the functions they contain.

[0094] The systems, components, and / or processes described herein may be implemented in hardware or a combination of hardware and software, centrally in a single processing system, or distributed across several interconnected processing systems with different elements. Any type of processing system or other device adapted to perform the methods described herein is suitable. A typical combination of hardware and software may be a processing system having computer-readable program code that controls the processing system to perform the methods described herein when loaded and executed. The systems, components, and / or processes may also be embedded in computer-readable storage devices, such as machine-readable computer program products or other data program storage devices, to explicitly embody a program of machine-executable instructions for performing the methods and processes described herein. These elements may also be embedded in application products that have features enabling the implementation of the methods described herein and can perform these methods when loaded into a processing system.

[0095] Furthermore, the configurations described herein may take the form of computer program products embodied in, for example, one or more computer-readable media having embodied computer-readable program code, stored therein. Any combination of one or more computer-readable media may be used. The computer-readable media may be computer-readable signal media or computer-readable storage media. The phrase "computer-readable storage medium" means a non-temporary storage medium. The computer-readable storage medium may, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or equipment, or any suitable combination thereof. More specific examples (not an exhaustive list) of computer-readable storage media would include: portable computer diskettes, HDDs (hard disk drives), SSDs (solid-state drives), ROMs, EPROMs or flash memory, portable CD-ROMs (compact disc read-only memory), DVDs (digital versatile discs), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing. In the context of this specification, computer-readable storage media may be any tangible medium that contains or can store programs for use by or in connection with an instruction execution system, device or apparatus.

[0096] In general, modules as used herein include routines, programs, objects, components, data structures, etc., that perform a specific task or implement a specific data type. In a further embodiment, memory generally stores the modules mentioned. Memory associated with a module may be a buffer or cache embedded within a processor, RAM, ROM, flash memory, or another suitable electronic storage medium. In a further embodiment, modules as envisioned by the present disclosure may be implemented as ASICs, as hardware components of a system on a chip (SoC), as a programmable logic array (PLA), or as another suitable hardware component embedded in a defined set of configurations (e.g., instructions) for performing the disclosed functions.

[0097] Program code embodied on a computer-readable storage medium may be transmitted using any suitable medium, including but not limited to wireless, wired, optical fiber, cable, RF (radio frequency), or any suitable combination thereof. Computer program code for performing operations for the current configuration may be written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Java®, Smalltalk, C++, or similar, and conventional procedural programming languages ​​such as the C programming language or similar. Program code may run entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer via any type of network, including a LAN (local area network) or a WAN (wide area network), or the connection may be made to an external computer (e.g., via the Internet using an Internet Service Provider).

[0098] As used herein, the terms “a” and “an” are defined as one or more. As used herein, the term “plurality” is defined as two or more. As used herein, the term “another” is defined as at least two or more. As used herein, the terms “including” and / or “having” are defined as “comprising” (i.e., open language). As used herein, the phrase “at least one of … and …” refers to and exhaustively includes one or more relatedly listed items and any possible combinations thereof. For example, the phrase “at least one of A, B, and C” includes A, B, C or any combination thereof (e.g., AB, AC, BC, or ABC).

[0099] For the sake of simplicity and clarity of explanation, it will be understood that, where appropriate, reference numerals are repeated between different drawings to point to corresponding or similar elements. In addition, the description outlines numerous characteristic details to provide a thorough understanding of the embodiments described herein. However, those skilled in the art will understand that the embodiments described herein may be put into practice using various combinations of these elements.

[0100] Embodiments of this specification may be embodied in other forms without departing from their spirit or essential characteristics. Therefore, the following claims, rather than the prior specifications, should be used to refer to this scope.

Claims

1. The processor has memory that is communicatively coupled to it. The aforementioned memory is An acquisition module, which includes instructions to cause the processor to acquire image data from the detector, which uses a lens to change multiple angles of light for each area of ​​the detector, when executed by the processor, When executed by the aforementioned processor, the processor: The kernel is mapped to the image data according to the size of the view and kernel associated with the area. A decision module that includes instructions for processing the image data using an ML (machine learning) model to create depth according to the size of the kernel, and stores the following: The determination module uses the ML model to classify objects in the scene related to the image data, A camera system further comprising instructions for generating a distribution of spatial points including the object related to the depth.

2. The camera system according to claim 1, wherein the lens has a gradient that becomes gentler or is configured to change the plurality of angles of the light related to the area of ​​the detector.

3. The camera system according to claim 1, wherein the lens includes one or more filter elements for directing the light to change the plurality of angles for each pixel or quadrant of the detector.

4. The camera system according to claim 1, wherein the image data includes a plurality of overlapping views from various angles that vary according to a parameter defined for depth in relation to any one of the refraction, filtering, and orientation of the light.

5. The camera system according to claim 1, wherein the acquisition module includes commands for acquiring the image data, and further includes commands for processing the light using a resonant waveguide grating (RWG) operably connected to the lens, and for changing the plurality of angles of the light.

6. The camera system according to claim 5, wherein the acquisition module further includes instructions for transmitting the light to the area of ​​the detector at a predetermined angle and wavelength according to the bandwidth of the RWG.

7. The camera system according to claim 1, wherein the area corresponds to one or more pixels of the detector associated with the plurality of angles.

8. When executed by a processor, the processor Image data is acquired from the detector, which uses lenses to change multiple angles of light for each area of ​​the detector, according to a reference. The kernel is mapped to the image data according to the size of the view and kernel associated with the area. The image data is processed using an ML (machine learning) model to create depth according to the size of the kernel. Using the aforementioned ML model, objects in the scene related to the image data are classified. A non-temporary, computer-readable medium containing instructions for generating a distribution of spatial points including the object related to the depth.

9. The non-transient computer-readable medium according to claim 8, wherein the lens has a gradient that becomes gentler or is configured to change the plurality of angles of the light related to the area of ​​the detector.

10. The non-transient computer-readable medium according to claim 8, wherein the lens includes one or more filter elements for directing the light to change the plurality of angles for each pixel or quadrant of the detector.

11. The non-temporary computer-readable medium according to claim 8, wherein the image data includes a plurality of overlapping views from various angles that vary according to a parameter defined for depth in relation to any one of the refraction, filtering, and orientation of the light.

12. Image data is acquired from the detector according to a reference, using a lens to change multiple angles of light for each area of ​​the detector, Mapping the kernel to the image data according to the size of the view and kernel associated with the area, The image data is processed using a machine learning (ML) model to create depth according to the size of the kernel, Using the aforementioned ML model, classify objects in the scene related to the image data, To generate a distribution of spatial points including the object related to the depth, Methods that include...

13. The method according to claim 12, wherein the lens has a gradient that becomes gentler or is configured to change the plurality of angles of the light related to the area of ​​the detector.

14. The method according to claim 12, wherein the lens includes one or more filter elements for directing the light to change the plurality of angles by the pixels or quadrants of the detector.

15. The method according to claim 12, wherein the image data includes a plurality of overlapping views from various angles that vary according to a parameter defined for depth in relation to any one of the refraction, filtering, and orientation of the light.

16. The method according to claim 12, further comprising obtaining the image data by processing the light using a resonant waveguide grating (RWG) operably connected to the lens to change the plurality of angles of the light.

17. The method according to claim 16, further comprising transmitting the light to the area of ​​the detector at a predetermined angle and wavelength according to the bandwidth of the RWG.

18. The method according to claim 12, wherein the area corresponds to one or more pixels of the detector associated with the plurality of angles.