Array camera for extracting 2D and 3D information from a scene
The array camera system with varied optical components and image processing techniques effectively captures and enhances two-dimensional and three-dimensional scene information, addressing the limitations of existing devices by leveraging blur and parallax metrics for depth extraction.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- NIL TECH APS (DK)
- Filing Date
- 2024-04-19
- Publication Date
- 2026-04-21
AI Technical Summary
Existing imaging devices struggle to efficiently capture both two-dimensional and three-dimensional information from a scene, particularly due to limitations in optical characteristics and focal lengths of lenses, which hinder the extraction of high-resolution and depth information.
An array camera system with multiple imaging components, each having distinct optical characteristics such as different focal lengths and F-numbers, combined with image processing techniques to extract three-dimensional information by analyzing blur and parallax metrics, utilizing machine learning networks trained on depth information.
Enables the extraction of high-resolution two-dimensional and three-dimensional information with reduced optical component costs and footprint, allowing for applications like depth enhancement in images and reduced manufacturing costs through planar wafer-level processes.
Smart Images

Figure 2026512906000001_ABST
Abstract
Description
Technical Field
[0001] Background This specification relates to imaging devices such as array cameras. An array camera can be based on an array of lenses. Individual images produced by each lens in the array can, in some cases, be combined to produce an image having a higher resolution than the individual images.
Summary of the Invention
Means for Solving the Problems
[0002] Overview This specification relates to array cameras and, in particular, describes techniques for array cameras that can be used to obtain two-dimensional and / or three-dimensional information (e.g., depth information) about a scene recorded in an image.
[0003] Generally, one or more aspects of the subject matter described in this specification can be embodied in an imaging device. The imaging device includes at least one sensor including a plurality of photosensitive regions operable to capture respective images of a scene being imaged, a first imaging component, and a second imaging component. Each of the first and second imaging components can be individually configured to capture an image of the scene to extract three-dimensional information of the scene. The first imaging component can have a first optical characteristic, the second imaging component can have a second optical characteristic, and the difference between the first optical characteristic and the second optical characteristic is higher than a threshold value of a predetermined difference in optical characteristics.
[0004] An implementation of the embodiment may include one or more features. A first optical feature may relate to a first focal length, a second optical feature may relate to a second focal length, and the difference between the first and second focal lengths is greater than a predetermined focal length difference. A first optical feature may relate to a first F-number, a second optical feature may relate to a second F-number, and the difference between the first and second F-numbers is greater than a predetermined focal length difference.
[0005] The imaging device may include an image processor that is operable to receive signals indicating each image captured by the photosensitive area, and the image processor may be further operable to extract three-dimensional information of the scene based on the difference between a first optical feature and a second optical feature.
[0006] In general, one or more aspects of the subject matter described herein can be embodied by one or more methods (and further, one or more non-temporary computer-readable media that specifically encode computer programs operable to cause one or more processors to perform operations). The method includes obtaining a comparative index of a first image and a second image of a scene, wherein the first image is captured by a first image component and the second image is captured by a second image component. The first imaging component may have a first optical feature, and the second imaging component may have a second optical feature, wherein the difference between the first and second optical features is greater than a predetermined threshold of difference between optical features.
[0007] One or more aspects of the subject matter described herein may also be embodied in one or more systems, which include one or more processors and a computer-readable medium storing instructions for causing the one or more processors to perform an operation, the operation including obtaining a comparative index of a first image and a second image of a scene, the first image being captured by a first image component and the second image being captured by a second image component. The first imaging component may have a first optical feature, and the second imaging component may have a second optical feature, the difference between the first optical feature and the second optical feature being greater than a predetermined threshold of difference between optical features.
[0008] The first optical feature can be related to the first focal length, the second optical feature can be related to the second focal length, and the difference between the first and second focal lengths is greater than a predetermined difference in focal length. The first optical feature can be related to the first F-number, the second optical feature can be related to the second F-number, and the difference between the first and second focal lengths is greater than a predetermined difference in focal length.
[0009] Obtaining comparative metrics may include obtaining blur comparison metrics and / or parallax metrics. Extracting three-dimensional information from a scene based on comparative metrics may include extracting three-dimensional information based on blur comparison metrics and / or parallax metrics.
[0010] Three-dimensional information can be extracted by a machine learning network trained to extract three-dimensional information based on a blur comparison index and / or a disparity index. The machine learning network can be trained based on a deep learning algorithm.
[0011] Certain embodiments of the subject matter described herein may be implemented to achieve one or more of the following advantages. The systems and techniques described herein may be used to provide a snapshot parallel imaging system using a one-dimensional (1D) or two-dimensional (2D) optical array that projects light onto a sensor or an array of sensors. Post-processing algorithms may be used to compute additional information about the captured scene. Among other applications, post-processing may be used to add depth information to existing images (e.g., RGB images), to restore high-resolution information from low-resolution inputs (e.g., super-resolution, refocus, virtual viewpoint, etc.), and / or to reconstruct high dynamic range (HDR). Furthermore, high performance in computing additional features (such as depth) can be achieved while at the same time the cost and footprint of the optical components and computing units can be optimized. Array cameras including meta-optical elements can offer several advantages in terms of reduced cost and small footprint thanks to planar wafer-level manufacturing processes. Furthermore, each channel can be independently controlled and tuned with respect to different functions, wavelengths, and / or polarizations of light. This makes the meta-optical elements particularly suitable for array cameras and parallel imaging systems.
[0012] Details of one or more embodiments of the subject matter described herein are described in the accompanying drawings and the following description. Other features, aspects, and advantages of the present invention will become apparent from the description, drawings, and claims. [Brief explanation of the drawing]
[0013] [Figure 1] This figure shows an example of an array camera that can generate images of a scene and operate to extract 3D information from that scene. [Figure 2] This figure shows an example of an array camera that can generate images of a scene and operate to extract 3D information from that scene. [Figure 3]This figure shows an example of an array camera that can generate images of a scene and operate to extract 3D information from that scene. [Figure 4] This figure shows an example of an array camera that incorporates folded optics and can be operated to generate images of a scene and extract 3D information from the scene. [Figure 5] Figures 1-3 show flowcharts illustrating the operation of one of the imaging devices. [Figure 6] This is a schematic diagram of a data processing system, including a data processing device that can be programmed as a client or a server, and the implementation techniques described in this document. [Modes for carrying out the invention]
[0014] Similar reference numbers and names in various figures indicate similar elements. Detailed explanation Figures 1-3 show examples of imaging devices that can be operated to generate images of a scene and extract 3D information from the scene. Figure 1 shows an example of an imaging device 100 (e.g., an array camera) which includes two imaging components, such as metalens 105A and 105B, along with apertures 110A and 110B, respectively, which can direct light to the photosensitive parts 120A and 120B of an imaging sensor to capture images 102A and 102B, respectively. Lenses 105A and 105B may have different optical characteristics. In the example of Figure 1, lenses 105A and 105B have different focal lengths. For example, the difference in focal lengths of lenses 105A and 105B may be higher than a predetermined threshold. The resulting difference in optical characteristics enables the extraction of 2D and 3D information 170. Lenses 105A and 105B may be designed such that one lens provides a sharper 2D image than the other. For example, at least one lens (e.g., lens 105A) provides a sharp 2D image. In the example in Figure 1, the focal length of lens 105A matches the distance to the corresponding photosensitive area 120A, producing a sharp image 102A. In the example in Figure 1, the focal length of lens 105B does not match the distance to the corresponding photosensitive area 120B, producing a blurred image 102B. For example, 3D information can be extracted by the image processing system 160 from the fact that subjects at different distances from the lens are captured with different amounts of blur in the image. During operation, a comparison index based on the amount of blur in images 102A and 102B captured by lenses 105A and 105B can be used to extract 3D information. In addition, or instead, parallax (e.g., angular parallax) caused by the different positions of lenses 120A and 102B can also be used by the imaging processing system 160 to extract 3D information. The resulting 3D information can be used for applications such as eye tracking and facial recognition. In some examples, additional imaging components, each with additional lenses having focal lengths between the lowest and highest focal lengths in the array camera, can be added to improve the depth map.
[0015] Figure 2 shows an example of an imaging device 200 (e.g., an array camera) that includes two imaging components, such as metalens 205A and 205B, which can direct light to the photosensitive parts 220A and 220B of an imaging sensor to capture images, respectively. The imaging components 105A and 105B may have different optical characteristics. In the example in Figure 2, the imaging components of lenses 205A and 205B have different f-numbers. For example, the difference in f-numbers between lenses 205A and 205B may be higher than a predetermined threshold. For example, apertures 210A and 210B may be selected so that one of the lenses produces a sharper image than the other. For example, one imaging component with a very low f-number may be used to greatly change the blur at a large subject distance (e.g., up to at least 0.5m) (e.g., approximately 1.0). The other imaging component may have a very large f-number (e.g., approximately 1.5 to 2.5). For example, 3D information (e.g., depth) can be extracted by an image processing system, such as the one shown in Figure 1, from the fact that subjects at different distances from the lens have different amounts of blur. A comparison index based on the amount of blur in the images captured by lenses 205A and 205B can be used for extracting 3D information. In addition, or instead, the parallax (e.g., angular parallax) caused by the different positions of lenses 205A and 205B can also be used by the image processing system to extract 3D information. In some examples, additional imaging components, each with additional lenses having f-numbers covering the region between the lowest and highest f-numbers, can be added to improve the depth map.
[0016] Figure 3 shows an example of an imaging device 300 (e.g., an array camera) that includes two imaging components, such as metalens 305A and 305B, which can direct light to the photosensitive parts 220A and 220B of an imaging sensor to capture images, respectively. In the example in Figure 3, lenses 305A and 305B have different focal lengths. Furthermore, apertures 310A and 310B are selected such that one of the lenses produces a sharper image than the other. The resulting differences in optical characteristics allow for the extraction of 2D and 3D information. For example, 3D information can be extracted by an image processing system, such as the image processing system shown in Figure 1, from the fact that subjects at different distances from the lens have different amounts of blur. A comparative index based on the amount of blur in the images captured by lenses 205A and 205B can be used to extract 3D information. In addition, or instead, the parallax (e.g., angular parallax) caused by the different positions of lenses 205A and 205B can also be used to extract 3D information.
[0017] In some examples, additional imaging components may be added to improve the depth map, each comprising additional lenses having a focal length between the minimum and maximum focal lengths of the array camera, and / or an F-number covering a region between the minimum and maximum F-numbers.
[0018] Figure 4 shows an example of an imaging device 400 (e.g., an array camera) that incorporates a bent optical system and is capable of generating an image of a scene and extracting 3D information from the scene. In a camera module, the optical Z height refers to the distance from the photoactive surface of the image sensor to the outermost point of the lens. This distance is often called the total track length (TTL). Reducing the optical Z height or TTL is often desirable to achieve low-profile camera modules, which can facilitate the integration of the camera module into compact electronic or other devices. Since many applications require low TTL, array cameras can be combined with a bent optical system to maintain low TTL. The imaging device 400 includes two imaging components, such as metalens 405A and 405B, along with apertures 410A and 410B, respectively, which can direct light to the photosensitive parts 420A and 420B of the image sensor to capture their respective images. Furthermore, mirrors 415A and 415B may be used to reduce TTL. Similar to the imaging devices in Figures 1-3, information can be extracted by an image processing system, such as the one shown in Figure 1, from the fact that subjects at different distances from the lens have different amounts of blur. A comparative index based on the amount of blur in the images captured by lenses 405A and 405B can be used for 3D information extraction. In addition, or instead, the parallax (e.g., angular parallax) caused by the different positions of lenses 405A and 405B can also be used by the imaging processing system to extract 3D information.
[0019] Additional imaging components may be added to improve the depth map, each comprising additional lenses having a focal length between the minimum and maximum focal lengths of the array camera, and / or an F-number covering a region between the minimum and maximum F-numbers.
[0020] The imaging devices shown in Figures 1-4 include two lenses, but the imaging device may include multiple lenses, such as an array camera.
[0021] As shown in FIG. 5, in operation, the imaging devices of FIGS. 1-4 can be used to extract 2D and / or 3D information. For example, 3D information can be extracted by an image processing system such as image processing system 160 from the fact that objects at different distances from the lens are captured with different amounts of blur in the image. The relationship between the distance from the lens and the amount of blur can be determined from one or more images depicting known distances. As soon as the relationship is determined, this can be used to infer an unknown distance from the amount of blur measured in the captured image. Additionally or alternatively, parallax (e.g., angular parallax) caused by different positions of the lenses in the array can also be used by image processing system 160 to extract the 3D information. In some examples, image processing system 160 can include a trained machine learning network to process comparison metrics such as blur metrics and parallax metrics to determine three-dimensional information such as depth. In some cases, image processing system 160 can incorporate artificial intelligence and may include iterative and / or deep learning techniques. (In FIG. 5, 510). The machine learning network can be trained with a set of images of known 2D / 3D information and known camera parameters that can be associated during training to input blur amounts and / or parallax metrics.
[0022] [[ID=⑥]]Once a trained machine learning network is obtained, comparison metrics of the first and second images captured by the first and second imaging components having different optical feature lengths and / or F-values (e.g., optical feature lengths and / or F-value differences higher than a threshold) can be obtained at 520. The comparison metrics can be blur amounts and / or parallax metrics.
[0023] At 530, two-dimensional and / or three-dimensional information can be extracted based on comparison metrics such as blur and / or parallax comparison metrics. For example, the obtained and trained machine learning network can be used to process the comparison metrics as input and obtain corresponding two / three-dimensional information such as depth information.
[0024] FIG. 6 is a schematic diagram of a data processing system including a data processing apparatus 500 that can be programmed as a client or as a server. The data processing apparatus 500 is connected to one or more computers 690 via a network 680. Only one computer is shown in FIG. 6 as the data processing apparatus 500, but multiple computers can be used. The data processing apparatus 500 includes various software modules that can be distributed between the application layer and the operating system. These can include executable and / or interpretable software programs or libraries, including tools and services of one or more programs 604 for image processing. The program 604 can execute one or more machine learning methods for 2D / 3D information extraction. Further, the program 604 can potentially execute manufacturing control operations of components of an imaging device (e.g., generation and / or application of specifications for performing the manufacture of meta-optical elements for an array camera). The number of software modules used can vary from one implementation to another. Further, the software modules can be distributed in one or more data processing apparatuses connected by one or more computer networks or other suitable communication networks.
[0025] The data processing unit 500 also includes hardware or firmware equipment, including one or more processors 612, one or more additional devices 614, a computer-readable medium 616, a communication interface 618, and one or more user interface devices 620. Each processor 612 is capable of processing instructions for execution within the data processing unit 500. In some implementations, the processors 612 are single or multi-threaded processors. Each processor 612 is capable of processing instructions stored in a storage device, such as the computer-readable medium 616 or one of the additional devices 614. The data processing unit 500 uses the communication interface 620 to communicate with one or more computers 690, for example, via a network 680. Examples of user interface devices 620 include displays, cameras, speakers, microphones, haptic feedback devices, keyboards, mice, and VR and / or AR equipment. The data processing device 500 can store instructions for performing operations related to the above-mentioned program, for example, on a computer-readable medium 616, or on one or more additional devices 614, such as a hard disk device, an optical disk device, a tape device, and a solid-state storage device.
[0026] The subject matter and functional operating embodiments described herein may be implemented in digital electronic circuits, or in computer software, firmware, or hardware, or a combination thereof, including the structures disclosed herein and their structural equivalents. Embodiments of the subject matter described herein may be implemented using one or more modules of computer program instructions encoded on a non-temporary computer-readable medium for execution by a data processing device or for control of the operation of a data processing device. The computer-readable medium may be a manufactured product, such as a hard drive in a computer system, an optical disc sold through a retail channel, or an embedded system. The computer-readable medium may be acquired separately and subsequently encoded in one or more modules after the delivery of one or more modules of computer program instructions, for example, via a wired or wireless network. The computer-readable medium may be a machine-readable storage device, a machine-readable storage substrate, a memory device, or a combination thereof.
[0027] The term "data processing device" encompasses all devices, machines, and equipment for processing data, including, for example, programmable processors, computers, or multiple processors or computers. In addition to hardware, a device may include code that creates the execution environment for the computer program in question, such as processor firmware, protocol stacks, database management systems, operating systems, runtime environments, or a combination of one or more of these. Furthermore, a device may employ a variety of different computing model infrastructures, such as web services, distributed computing, and grid computing infrastructures.
[0028] Computer programs (also commonly known as programs, software, software applications, scripts, or code) can be written in any suitable form of programming language, including compiled or interpreted languages, declarative or procedural languages, and can be deployed in any suitable form, including as standalone programs or modules, components, subroutines, or other units suitable for use in a computing environment. Computer programs do not necessarily correspond to files in a file system. A program may be stored in a single file dedicated to the program in question, or in multiple collaborative files (e.g., a file storing one or more modules, subprograms, or parts of code), or as part of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document). Computer programs may be employed to run on a single computer, or on multiple computers distributed across one or more locations and interconnected by a communication network.
[0029] The processes and logic flows described herein may be executed by one or more programmable processors that execute one or more computer programs to perform a function by acting on input data and producing an output. The processes and logic flows may also be executed by logic circuits for specific purposes, such as FPGAs (Field-Programmable Gate Arrays) or ASICs (Application-Specific Integrated Circuits), and the devices may also be implemented as logic circuits for specific purposes.
[0030] A suitable processor for executing a computer program includes, for example, both general-purpose and dedicated microprocessors, and one or more processors in any type of digital computer. Generally, a processor will receive instructions and data from read-only memory, random-access memory, or both. Essential elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Generally, a computer will also include one or more mass storage devices for storing data, such as magnetic disks, magneto-optical disks, or optical disks, or will be operationally coupled to receive data from or transfer data to such devices, or both. However, a computer is not required to have such devices. Furthermore, a computer may be incorporated into other devices, for example, mobile phones, personal digital assistants (PDAs), portable audio or video players, game consoles, Global Positioning System (GPS) receivers, or portable storage devices (e.g., Universal Serial Bus (USB) flash drives). Devices suitable for storing computer program instructions and data include, for example, semiconductor memory devices such as EPROM (Erasable Programmable Read-Only Memory), EEPROM (Electronically Erasable Programmable Read-Only Memory), and flash memory devices; magnetic disks, such as internal hard disks or removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks, encompassing all forms of non-volatile memory, media, and memory devices. Processors and memory may be supplemented by or integrated into dedicated logic circuits.
[0031] To provide user interaction, embodiments of the subject matter described herein may be implemented in a computer having a display device, such as an LCD (liquid crystal display) display device, an OLED (organic light-emitting diode) display device, or other monitor for displaying information to the user, as well as a keyboard and pointing device, such as a mouse or trackball, to which the user can provide input to the computer. Other types of devices may also be used to provide user interaction; for example, the feedback provided to the user may be any suitable form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback; and input from the user may be received in any suitable form, including acoustic, speech, or tactile input.
[0032] A computing system may include clients and servers. Clients and servers are generally geographically distant from each other and typically interact via a communication network. The client-server relationship arises thanks to computer programs running on each computer and having a client-server relationship with each other. Embodiments of the subject matter described herein may be implemented in a computing system or a combination of one or more such backend, middleware, or frontend components, including, for example, a data server, a backend component, or a middleware component, such as an application server, or a frontend component, such as a client computer having, for example, a graphical user interface or a browser user interface, with which a user can interact in the implementation of the subject matter described herein. The components of the system may be interconnected by any suitable form or medium of digital data communication, such as a communication network. Examples of communication networks include local area networks ("LANs") and wide area networks ("WANs"), internetworks ("the Internet"), and peer-to-peer networks (e.g., ad-hoc peer-to-peer networks).
[0033] This specification includes many implementation details, which should not be construed as limitations on the scope of what is claimed or may be claimed, but rather as descriptions of features specific to particular embodiments of the disclosed subject matter. Explicit features described herein in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented separately in multiple embodiments or in any suitable subcombination. Furthermore, features may be described above as working in a combination, and may even be initially claimed as such, but one or more features may be removed from the claimed combination, and the claimed combination may be directed towards a subcombination or a variation of a subcombination.
[0034] Similarly, although the operations are depicted in a specific order in the diagrams, this should not be understood as requiring such operations to be performed in a specific order or sequentially shown, or that all depicted operations must be performed, in order to achieve the desired result. In certain circumstances, multitasking and parallel processing may be advantageous. Furthermore, the separation of various system components in the embodiments described above should not be understood as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together into a single software product or packaged into multiple software products.
[0035] Thus, specific embodiments of the invention have been described. Other embodiments are within the scope of the following claims. In addition, the operations described in the claims can be performed in different orders and still achieve the desired results.
Claims
1. An imaging device, A sensor comprising at least one sensor having multiple photosensitive areas, each area operable to capture a separate image of the scene being imaged, An imaging device comprising a first imaging component and a second imaging component, each of the first and second imaging components individually configured to capture an image of the scene in order to extract three-dimensional information of the scene, wherein the first imaging component has a first optical feature, and the second imaging component has a second optical feature, and the difference between the first optical feature and the second optical feature is higher than a predetermined threshold for the difference between optical features.
2. The imaging device according to claim 1, wherein the first optical feature is related to a first F-number of the first imaging component, the second optical feature is related to a second F-number of the second imaging component, and the difference between the first F-number and the second F-number is greater than a predetermined F-number difference.
3. The imaging device according to claim 1 or claim 2, wherein the first optical feature is related to a first focal length, the second optical feature is related to a second focal length, and the difference between the first focal length and the second focal length is greater than a predetermined focal length difference.
4. The device according to any of the above claims, comprising an image processor operable to receive signals indicating each of the images captured by the photosensitive region, wherein the image processor is further operable to extract three-dimensional information of the scene based on the difference between the first optical feature and the second optical feature.
5. A method performed by an image processor, The system includes obtaining a comparison index between a first image and a second image of a scene, wherein the first image is captured by a first imaging component, the second image is captured by a second imaging component, the first imaging component has a first optical feature, the second imaging component has a second optical feature, and the difference between the first optical feature and the second optical feature is higher than a predetermined threshold for the difference in optical features. A method comprising extracting three-dimensional information from the scene based on the aforementioned comparison index.
6. The method according to claim 5, wherein obtaining the comparison index comprises obtaining a blur amount comparison index and / or a parallax index, and extracting three-dimensional information from the scene based on the comparison index comprises extracting the three-dimensional information based on the blur comparison index and / or the parallax index.
7. The method according to claim 6, wherein the three-dimensional information is extracted by a machine learning network trained to extract the three-dimensional information based on the blur comparison index and / or the disparity index.
8. The method according to claim 7, wherein the machine learning network is trained based on a deep learning algorithm.