Automatic sensor pixel placement method optimized for multiple camera tasks
The system optimizes pixel alignment on image sensors for multiple camera tasks using a trainable pixel alignment layer and neural networks, addressing the challenge of balancing performance across diverse camera functions for improved image quality and functionality.
Patent Information
- Application Number
- JP2024549553
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2022-06-16
- Filing Date
- 2023-01-11
- Publication Date
- 2026-03-09
- Estimated Expiration
- 2043-01-11
AI Technical Summary
Existing methods fail to optimize the alignment of different pixel types on an image sensor for multiple camera tasks simultaneously, leading to suboptimal performance in balancing user priorities across various camera functions.
A system and method that utilizes a trainable pixel alignment layer and a stack of neural networks to optimize the alignment of multiple pixel types on an image sensor, taking into account user-determined priorities for multiple camera tasks, employing techniques like continuous relaxation and reinforcement learning to achieve a balanced performance across these tasks.
Enables automatic optimization of pixel alignment for improved performance in multiple camera tasks, balancing user-defined priorities through end-to-end training and differentiable approximations, enhancing image quality and functionality.
Smart Images

Figure 0007826503000001 
Figure 0007826503000002 
Figure 0007826503000003
Abstract
Description
[Technical Field]
[0001] (CROSS-REFERENCE TO RELATED APPLICATIONS) This application claims the benefit of U.S. Patent Application No. 17 / 952,077 (SYP347918US02), filed June 16, 2022, entitled "AUTOMATIC SENSOR SPARSE PLACEMENT METHOD OPTIMIZED FOR MULTIPLE CAMERA TASK," and U.S. Provisional Patent Application No. 63 / 268 / 386 (SYP347918US01), filed February 23, 2022, entitled "METHOD FOR PLACING SENSOR PIXELS," both of which are hereby incorporated by reference for all purposes as if fully set forth herein. [Background technology]
[0002] Modern camera image sensors include different types of pixels. Some types specialize in how they address aspects of the captured image quality, such as accurate color representation (using pixels of a color filter array) or high dynamic range (using highly sensitive pixels combined with a high threshold saturation limit). Other types do not determine the quality of the image captured by the sensor array, but are designed to address aspects related to downstream camera tasks, including image processing such as autofocus, object recognition, object removal, or depth determination. In these cases, various subtypes of phase-sensitive pixels are often relevant.
[0003] The pattern in which different types of pixels are arranged (aligned) across the image sensor surface plays a major role in determining the appearance of the captured image, some of which are directly related to the quality of the image captured by the image sensor, while others are related to the downstream camera task that provides the captured image as input for analysis. However, a particular pattern that is advantageous for performing one camera task of interest to a user may be detrimental to another camera task of interest to the user, and therefore it may be desirable to find an "optimized" pattern that balances the camera user's priorities across two or more tasks.
[0004] Manual design of an optimized pixel alignment pattern that results in good performance for one or more camera tasks that are considered high priority, at the expense of reduced performance for one or more camera tasks that are lower priority, is a time-consuming and expensive process for camera manufacturing.
[0005] Several approaches have been proposed to automate pixel alignment design with the sole goal of accurate color representation in the image output by the camera. In this case, the task is to maximize the fidelity of colors captured by the image sensor array, which is determined in large part by three or more subtypes of color filter array (CFA) pixels on the sensor array. All pixels are of the same basic type, responding to the luminance or intensity of incident light, and their outputs are simply combined to produce an image that shows the spatial variation of light emitted from the surface of the imaged object. Pixel subtypes differ from each other only in the spectral transmittance of filters fabricated on or above the top layer of pixels of a given subtype, which selectively block light outside the corresponding spectral band. In an RGB image sensor system, for example, three subtypes of CFA pixels are created by red, blue, and green filters, respectively. A typical alignment or arrangement of the filtered pixels is a 2x2 Bayer block pattern, repeated across the entire sensor array, with green filtered pixels located on the diagonal and red and blue filtered pixels occupying the other two spaces. It has been found that improved color performance can be achieved by a more complex alignment, in one example involving 4x2 pixel blocks of four different color filtered subtypes, although many other combinations can (and are) anticipated.
[0006] Another example of an automated approach to pixel alignment is the design process for a pattern of shutter functions that is superimposed on an image sensor's pixel array, in this case with the goal of maximizing the dynamic range of the captured image. The pixels in the array are of one basic type, capturing gray-level intensities (because there are no color filters in this case), and the "alignment pattern" is not actually a pattern of different pixel types or subtypes, but rather a spatial distribution on the sensor surface of a set of circuit parameters that determine exposure times that vary continuously across the array. The spatial distribution of exposure times across the pixel array primarily determines the achievable dynamic range in the images captured by those pixels.
[0007] To date, approaches have focused on the problem of optimizing the alignment of pixels on an image sensor for a single camera task, where only one type of pixel is relevant and determines the image quality aspect. No approach appears to address the problem of simultaneously optimizing the performance of multiple camera tasks, taking into account their relative priorities. In these situations, a mixture of different pixel types may be relevant, where one type may not be relevant to the perceived quality of the captured image, which can be of great practical importance to the end user; phase-sensitive pixels are one such example, required for such tasks as camera autofocus, 3D reconstruction, etc.
[0008] Thus, there is a need for a system and method that automates the design of pixel alignment for an image sensor having at least two different types of pixels in order to optimize overall camera performance in a manner that achieves a desired balance between the performance of two or more different tasks. Summary of the Invention
[0009] The present invention includes methods and systems for optimization of the alignment of X different types of pixels on an image sensor array for performance of N different camera tasks, where X and N are integers greater than one.
[0010] In one embodiment, the method includes obtaining an output of each type of pixel in a sensor for an input image, evaluating the quality of the output for each of the N camera tasks, and obtaining an optimal pixel alignment pattern by adjusting potential pixel alignment patterns with respect to performance of the N camera tasks according to a balance of user-determined priorities among the N camera tasks.
[0011] In another embodiment, a method includes simulating pixel responses to the training images, outputting X sensor images for each training image, where each sensor image arbitrarily corresponds to only one of X different types located at each possible location of the image sensor array; generating a pixel alignment pattern according to which the X different types of pixels are distributed on the image sensor array as a 2D tessellation in the trainable pixel alignment; subsampling each of the X sensor images according to the pixel alignment pattern and outputting X corresponding subsampled images; and simultaneously training a pixel alignment layer and a stack of N neural networks, each neural network uniquely corresponding to only one of the N different camera tasks, to receive and process one or more of the X subsampled sensor images to make a discrete and distinct selection of the X different types of pixels, such that the pixel alignment pattern is optimized for performance of the N camera tasks targeted at a user-determined balance of priorities among the N camera tasks.
[0012] In another embodiment, the apparatus includes one or more processors and logic encoded on one or more non-transitory media for execution by the one or more processors, the logic, when executed, outputting, for each training image, X sensor images, where each sensor image uniquely corresponds to only one of X different types of pixels located at each possible location of the image sensor array, simulating pixel responses to the training images; generating pixel alignment patterns in a trainable pixel alignment layer where the X different types of pixels are distributed over the locations of the image sensor array as a 2D tessellation; and generating pixel alignment patterns for the X corresponding pixels. and simultaneously training the pixel alignment layer and a stack of N neural networks, each neural network uniquely corresponding to only one of the N different camera tasks, such that the pixel alignment pattern is optimized for performance of the N camera tasks according to a user-determined balance of priorities among the N camera tasks, receiving and processing the X subsampled sensor images to make a discrete and distinct selection of the X different types of pixels.
[0013] In yet another embodiment, a system includes a pixel response simulator, a trainable pixel alignment layer, a sub-sampler, and a stack of N trainable neural networks, each neural network uniquely corresponding to only one of the N camera tasks; the pixel response simulator is configured to operate on training images and, for each training image, send X sensor images to the sub-sampler, each sensor image uniquely corresponding to pixels of only one of X different types located at each possible location on the image sensor array; and the trainable pixel alignment layer is configured to arrange the X different types of pixels as a 2D tessellation at the locations on the image sensor array. The method further comprises distributing the pixel alignment patterns over the sensor images to generate pixel alignment patterns and sending them to a subsampler, the subsampler being configured to subsample each of the X sensor images according to the pixel alignment patterns and sending one or more subsampled output images to one or more neural networks for processing, the N neural networks and pixel alignment layers being trained simultaneously using inner and outer feedback loops, respectively, to make discrete and distinct selections of the X different types of pixels such that the pixel alignment patterns are optimized for performance of the N camera tasks according to a user-determined balance of priorities among the N camera tasks.
[0014] A further understanding of the nature and advantages of specific embodiments disclosed herein may be realized by reference to the remaining portions of the specification and the attached drawings. [Brief explanation of the drawings]
[0015] [Figure 1] 1 is a flowchart illustrating method steps according to some embodiments of the present invention. [Figure 2] 1 is a block diagram illustrating a system according to some embodiments of the present invention. [Figure 3]3 is a portion of a system according to the embodiment of FIG. 2. [Figure 4] This is how the combined performance value can be calculated in terms of task losses in some of the embodiments of FIG. [Figure 5] 1 is a block diagram illustrating a system according to some embodiments of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0016] Described herein are embodiments of systems and methods that optimize the alignment of pixels of two or more different types on an image sensor array for performance of two or more camera tasks, taking into account the priority level of the tasks.
[0017] 1 is a flowchart illustrating steps of a method 100 according to some embodiments of the present invention. In step 101, the output of each type of pixel on an image sensor array is obtained given an input image. In step 102, the quality of the output for each of N camera tasks is evaluated for potential pixel alignment patterns on the sensor array. In step 103, an optimal pixel alignment pattern is obtained by adjusting the potential pixel alignment patterns for the performance of the N camera tasks according to a user-determined balance of priorities among the N camera tasks.
[0018] In some embodiments of method 100, obtaining an optimal pixel alignment pattern includes evaluating all possible combinations of pixel alignments on the sensor array. In other embodiments, it includes using one or more rule-based algorithms. In some embodiments of method 100, obtaining an optimal pixel alignment pattern includes using machine learning. The machine learning may utilize trainable pixel alignment layers, in some cases involving successive relaxation techniques. Alternatively, the machine learning may utilize reinforcement learning techniques.
[0019] In some embodiments of method 100, the sensor array includes a tessellation of identical pixel blocks, each block including a first arrangement of X different types of pixels, and obtaining an optimal pixel alignment pattern includes obtaining an optimal arrangement of the X different types of pixels in a block. In one subset of these embodiments, evaluating all possible combinations of optimal alignments of pixels in each block includes evaluating all possible combinations of optimal alignments of pixels in each block. In another subset of these embodiments, obtaining an optimal arrangement of pixels in each block includes using one or more rule-based algorithms. In yet another subset of these embodiments, obtaining an optimal arrangement of pixels in each block includes using machine learning, which may utilize a trainable pixel alignment layer, and in some cases involving successive relaxation techniques or alternatively, the machine learning may utilize reinforcement learning techniques.
[0020] In another embodiment of method 100, the sensor array includes a tessellation of first and second blocks of pixels, each of the first blocks including a first arrangement of X different types of pixels, and each of the second blocks including a second arrangement of the X different types of pixels that is different from the first arrangement, and obtaining the optimal pixel alignment pattern includes obtaining an optimal arrangement of the X different types of pixels in the first blocks and the second blocks. Subsets of embodiments can readily be envisioned for the above various options regarding different approaches to obtaining the optimal pixel alignment pattern (e.g., evaluating all possible combinations, using a rule-based algorithm, etc.).
[0021] In some embodiments of method 100, each pixel type is characterized by one or more associated adjustable parameters, such as phase angle, filter wavelength range, etc. In some embodiments of method 100, step 101, in which the output of each type of pixel on the sensor is obtained, involves simulating the response of that type of pixel. A description of how this may be done is provided below with reference to system 200 illustrated in FIG. 2. In other embodiments, obtaining those outputs involves measuring the response of each type of pixel, rather than simulating it. This may be done by physically capturing an image using a camera and directly recording the pixel response.
[0022] 2 illustrates a system 200 according to some embodiments of the present invention, which essentially performs method 100 using a particular combination of selections for implementing steps 101, 102, and 103. Other systems may be expected to use other selections without departing from the spirit and scope of the present invention.
[0023] A pixel response simulator 202 receives training images 201 from a training image database (not shown), generates multiple X sensor images 203, one for each type of pixel whose response is to be simulated, and sends the images to a subsampler 206. A trainable pixel alignment layer module 204 generates a pixel alignment pattern 205 for the mix of pixel types and sends the pattern to the subsampler 206. The subsampler 206 operates on each one of the received sensor images 203 in turn, extracting data for the corresponding pixel type at each location indicated for that pixel type in the pattern 205, and then tessellating a small portion of the resulting simulated sensor image on the sensor array. In this manner, the subsampler 206 generates X subsampled images 207, one image for each of the X types of pixels.
[0024] A stack 208 of N trainable networks receives the subsampled images 207, with each individual network dedicated to simulating only one camera processing task receiving one or more of the subsampled images appropriate to the operation of that particular task. So, for example, network 208-1 dedicated to simulating camera processing task #1 receives one or more subsampled images appropriate to the operation of task #1, network 208-2 for task #1 receives either the subsampled image or images associated with the network for task #2, etc. As described below with respect to Figures 3 and 4, there may be some overlap between the subsampled images sent to different networks, depending on the nature of the corresponding camera tasks, but this is not necessary.
[0025] In some embodiments, each of the N neural networks may receive all X subsampled images and process some of them using information not directly related to that particular network's particular camera task but that provides relevant side information. Each network in stack 208 provides an output value reflecting how well it currently performs its corresponding task, with 208-1 providing a performance value 209-1 for its performance of task #1, 208-2 providing 109-2, and network 208-N providing 209-N for the Nth task. The individual performance values are input to module 210, which calculates a combined performance value (another word for "combined performance value" in optimization literature is "objective function") 212 that reflects how well all N tasks will perform for the current pixel alignment pattern 205 and the current version of each network's internal configuration. In some scenarios, these parameters may be provided to module 210 as separate inputs (not shown).
[0026] The combined performance value 212 is fed back from the output of module 210 via simultaneous inner and outer feedback loops. The inner loop includes path A between module 210 and stack 208, allowing each of networks 208-1 through 208-N to self-learn and optimize performance value 212. In some embodiments, this means minimizing the corresponding combined loss value. The outer loop includes path B between module 210 and pixel alignment layer 204, allowing layer 204 to learn and optimize pixel alignment pattern 205, thereby optimizing performance value 212.
[0027] Training of the system 200 continues according to the process described above for each of a series of training images available from the database.
[0028] Modules 204 and 208 are tunable, allowing them to be optimized using clues from combined performance value 212 .
[0029] One possible way to implement this optimization is by using stochastic gradient descent, The performance value is increased by calculating the gradient of the performance value with respect to the tunable parameters (neural network weights) and applying updates to the tunable parameters in the direction of the gradient. In the stochastic gradient descent framework, all modules between the tunable modules and the performance value are required to be differentiable.
[0030] Given these requirements, it can be recognized that a trainable alignment layer, such as layer 204 in Figure 2, must be defined by a differentiable function. The gradient of the performance value must somehow be computed for different pixel selections. However, because the pixel selections are discrete, the gradient is not defined.
[0031] To address this, layer 104 is implemented using a method that approximates the selection of discrete pixels in a differentiable manner.
[0032] The methods used in some embodiments of the present invention fall under the umbrella of "continuous relaxation," a technique used to address a general category of problems that require derivatives and / or gradients to be calculated from non-differentiable discrete or classifier functions by reformulating them into approximations using continuous variables.
[0033] While continuous relaxation is one approach to approximating a non-differentiable function with a differentiable function, there are other methods that can be considered for other embodiments.
[0034] By implementing 204 with a differentiable approximation, the present invention allows sensor designers to optimize sensors where discrete and categorical design choices can be made, in contrast to prior art techniques that limit the optimization space to only continuous ones.
[0035] Figure 3 illustrates a portion of a system of the type shown in Figure 2 for one example embodiment in which there are three different pixel types to place on the sensor and two different camera tasks of interest. In this embodiment, pixel response simulator 302 receives training image 301 and outputs three sensor images 303A, 303B, and 303C, where 303A shows what the sensor output would be if the sensor array only had "A" type pixels at every pixel location on the sensor array, and 303B and 303C do the same for "B" and "C" type pixels, respectively.
[0036] The pixel alignment layer 304 generates a pattern 305 in the form of 4x4 blocks, which is repeated as a 2D tessellation on the surface of the image sensor array (not shown). Each block has five each of "A" and "C" type pixels, and six of "B" type pixels, the latter positioned diagonally, although it should be understood that this is only one "candidate" for an optimized arrangement.
[0037] Subsampler 306 operates on sensor images 303A, 303B, and 303C to extract data from pixels of types A, B, and C in each block indicated for that pixel type in pattern 305, and then tessellates the resulting fractional portions of the simulated sensor image on the sensor array. The resulting output is three subsampled images 307A, 307B, and 308C.
[0038] In this exemplary system with two camera tasks of interest, stack 308 consists of two corresponding trainable neural networks 308-1 and 308-2, delivering outputs 309-1 and 309-2. In the case shown, because the performance of network 308-1's camera task is affected only by how type A pixels perform, but not by type B or C, only subsampled image 307A is fed into that network as input. Network 308-1's task may be, for example, an object recognition task, determined entirely by type A pixels. At the same time, the performance of the camera task corresponding to network 308-2 may be affected to a greater extent by how type B and type C pixels perform; therefore, network 308-2 may require both images 307B and 307C to be provided as input. Network 308-2's task may be, for example, an "image quality" task, providing images with a high dynamic range, which may require a mix of highly sensitive pixels and a high saturation threshold.
[0039] In some systems similar to that shown in FIG. 3 , all two types of pixels may be relevant to the performance of both objective tasks. In other embodiments not shown, at least one pixel type may be relevant to both camera tasks, etc. As the number of variables increases, naturally, more pixel types and / or more camera tasks are modeled. A key feature of all embodiments of the current invention is that the X pixel types are distinct from one another and the N camera tasks are distinct from one another. In some scenes, at least one task achieves an image quality goal determined by one or more types of pixels in one subcategory of simple image capture pixels, while another task achieves a goal that relies on distinct types of pixels, such as phase detection pixels. Producing images with a high dynamic range is an example of an image quality task, while processing the captured images to identify objects of a particular shape within the scene is an example of a very difficult type of camera or image processing task. Other tasks may include generating a depth map or creating a semantic understanding of a scene.
[0040] 3, the outputs of the two networks are performance values 309-1 and 309-2, respectively, which are then fed into modules (as shown) corresponding to modules 210 of system 200, which operate on those values to generate a combined performance value that reflects how well the set of tasks performed for the three pixel types of interest and for the current version of pixel alignment pattern 305 for the current internal configuration of neural networks 308-1 and 308-2. The combined performance value is fed back to the neural network and pixel alignment layer 304 of stack 308 to generate improved network parameters and pixel alignment patterns by processes well known in the art of machine learning for system parameter optimization.
[0041] 4 shows how performance values are calculated and processed in some embodiments of the invention, where the performance values are expressed as loss values that are inversely related to how well a task is performed. In the scenario shown, there are N neural networks in a stack 408, with the first network outputting a performance value 408-1 in the form of a task loss L1, and so on, one by one, up to the last Nth network, which outputs a task loss L N Next, in module 410, the kth term in the function where k varies from 1 to N is calculated as f (α k , L k, ) a combined performance value is calculated as a function of each performance value and a corresponding weighting parameter α for each task. In some embodiments, this function is such that the kth item is k ×L k , the function may be linearly dependent on the individual performance values. k, ) 2 The weighting parameters indicating the relative priorities of the N tasks can be either pre-defined inputs embedded in the system or inputs as adjustable parameters set by the user to train the system on performance.
[0042] Another option is that the performance values immediately expected for the training system of the present invention are restated as loss values inversely related to how well the task is performed, but the combined performance value is calculated as a function of each performance value and two or more task-specific weighting parameters. In the two parameter case, for example, the kth term of the function is, for example, [(α k ) 2 *L k ] + [β k * (L k ) 2] Many variables can be predicted, some of which may be associated with more than two weighting parameters.
[0043] While the embodiments described above and shown in the figures all involve the use of a stack of individual neural networks, the invention is not necessarily limited to this particular architecture. Figure 5 illustrates some other embodiments in which there is one large "super" neural network 508 that receives all X sensor images as input and outputs N values corresponding to each task. The other elements of system 500 operate in the same manner as system 200 shown in Figure 2, as follows:
[0044] The pixel response simulator 502 receives training images 501 from a training image database (not shown), generates a plurality of X sensor images 503, one for each type of pixel whose response is to be simulated, and sends the images to a subsampler 506. The trainable pixel alignment layer module 404 generates a pixel alignment pattern 505 for the mix of pixel types and sends the pattern to the subsampler 506. The subsampler 506 operates on each received sensor image 503 in turn, extracting data for its corresponding pixel type at each location indicated for that pixel type in the pattern 505, and then tessellating the resulting functional portion of the simulated sensor image on the sensor array. In this manner, the subsampler 506 generates X subsampled images 507, one image for each of the X types of pixels.
[0045] However, trainable network 508 receives all images 597 and provides N output values, one for each of the N camera tasks of interest, indicating how well the system currently performs its corresponding task. The individual performance values are input to module 510, which calculates a combined performance value 512 that reflects how well all of the N tasks are performing for the current version of pixel alignment pattern 505 and the current internal configuration of super-network 508. The calculation involves weighting the individual performance values by a predetermined task priority parameter.
[0046] Embodiments of the present invention offer significant advantages over the prior art in this area in providing a system and method for automatic optimization of pixel alignment of different underlying types for a camera intended to perform more than one task, taking into account user-determined priorities between those tasks. In some embodiments, this is achieved by performing end-to-end training of a system comprising a stack of task-specific neural networks and a trainable pixel alignment layer, with the user setting one or more parameters that balance competing aspects of the camera's performance.
[0047] Although descriptions are given with respect to specific embodiments herein, these specific embodiments are merely illustrative and not limiting.
[0048] Any suitable programming language may be used to implement the routines of a particular embodiment, such as C, C++, Java, assembly language, etc. Different programming techniques may be used, such as procedural and object-oriented. The routines may be executed on a single processing device or multiple processors. Although steps, operations, or computational uses may be given in a particular order, this order may be changed in different particular embodiments. In some particular embodiments, multiple steps shown in this specification as sequential may be performed simultaneously.
[0049] Certain embodiments may be implemented in a computer-readable storage medium for use by or in connection with an instruction execution system, apparatus, or device. Certain embodiments may be implemented in the form of control logic in software or hardware, or a combination of both. The control logic, when executed by one or more processors, may be operable to perform what is described in certain embodiments.
[0050] Certain embodiments may be implemented using a programmed general-purpose digital computer, using application-specific integrated circuits, programmable logic devices, field programmable gate arrays, optical, chemical, biological, quantum, or nano-optical systems, components, and mechanisms that can be used. In general, the functionality of certain embodiments may be achieved by any means known in the art. Distributed, networked systems, components, and / or circuits may be used. Communication or transfer of data may be by wire, wireless, or any other means.
[0051] It will be appreciated that one or more elements depicted in the drawings / figures may also be implemented in a more discrete or integrated manner, or in some cases removed or rendered inoperable, as may be useful in a particular application. It is also within the intent and scope to implement a program or code stored on a machine-readable medium that enables a computer to perform any of the methods described above.
[0052] A "processor" includes any suitable hardware and / or software system, mechanism, or component that processes data, signals, or other information. A processor may include a system with a general-purpose central processing unit, a microprocessor, dedicated circuitry for implementing a function, or other system. Processing need not be limited to a geographic location or have time limitations. For example, a processor may perform its functions in "real time," "offline," "batch mode," etc. Portions of processing may be performed at different times and in different locations by different (or the same) processing systems. Examples of processing systems may include servers, clients, end-user devices, routers, switches, networked storage devices, etc. A computer may be any processor in communication with a memory. Memory may be random access memory (RAM), a magnetic or optical disk, or any suitable processor-readable storage medium suitable for storing instructions for execution by a processor.
[0053] As used herein and throughout the claims that follow, the terms "a," "an," and "the" include plural references unless the context clearly dictates otherwise. Furthermore, as used herein and throughout the claims that follow, the meaning of "in" includes "in" and "on," unless the context clearly dictates otherwise.
[0054] Thus, while specific embodiments have been described herein, it will be recognized that latitude in modification, various changes, and substitutions are contemplated in the foregoing disclosure, and that in some cases, some features of a specific embodiment may be utilized without the corresponding use of other features without departing from the scope and intent set forth. Therefore, many modifications may be made to adapt a particular situation or material to the essential scope and spirit.
Claims
1. 1. A method for optimizing the alignment of X different types of pixels on a sensor array for a camera capable of performing N different camera tasks, where X and N are integers greater than 1, the method comprising: obtaining an output for each of the X types of pixels on the sensor array for an input image to the sensor array; assessing the quality of the output for each of the N different camera tasks; A method comprising: obtaining an optimal pixel alignment pattern by evaluating output values reflecting the degree of performance of each of the N different camera tasks determined by a user, taking into account the weighting in each camera task, and adjusting a pixel alignment pattern that results in a relatively high overall output value as a potential pixel alignment pattern that is a candidate for the optimal pixel alignment pattern for performance of the N different camera tasks according to a balance of priorities.
2. The method of claim 1 , wherein obtaining the optimal pixel alignment pattern comprises evaluating all possible combinations of pixel alignments on the sensor array.
3. The method of claim 1 , wherein obtaining the optimal pixel alignment pattern comprises using one or more rule-based algorithms.
4. The method of claim 1 , wherein obtaining the optimal pixel alignment pattern comprises using machine learning.
5. The method of claim 4 , wherein using machine learning comprises utilizing a trainable pixel alignment layer.
6. The method of claim 5 , wherein utilizing an alignment layer of trainable pixels comprises using a successive relaxation technique.
7. The method of claim 4 , wherein the machine learning comprises utilizing reinforcement learning techniques.
8. the sensor array is a tessellation of pixel blocks having the same alignment pattern of the X types of pixels, each block containing an arrangement of the X different types of pixels; Obtaining an optimal pixel alignment pattern includes obtaining an optimal arrangement of the X different types of pixels in a block. The method of claim 1.
9. The method of claim 8 , wherein obtaining the optimal placement of pixels in each block includes evaluating all possible combinations of alignments of pixels in the block.
10. The method of claim 8 , wherein the optimal pixel alignment pattern comprises using one or more rule-based algorithms.
11. The method of claim 8 , wherein obtaining the optimal pixel alignment pattern comprises using machine learning.
12. The method of claim 8 , wherein using machine learning comprises utilizing a trainable pixel alignment layer.
13. The method of claim 12 , wherein utilizing an alignment layer of trainable pixels comprises using a successive relaxation technique.
14. The method of claim 8 , wherein the machine learning comprises utilizing reinforcement learning techniques.
15. the sensor array is a tessellation of first and second blocks of pixels, each of the first blocks including a first arrangement of the X different types of pixels and each of the second blocks including a second arrangement of the X different types of pixels different from the first arrangement; obtaining the optimal pixel alignment pattern includes obtaining an optimal pixel arrangement pattern in the first block and the second block; The method of claim 1.
16. The method of claim 1 , wherein each pixel type is characterized by one or more adjustable parameters associated with said each pixel type.
17. The method of claim 1 , wherein obtaining the output of each type of pixel on the sensor array comprises simulating the response of that type of pixel.
18. The method of claim 1 , wherein obtaining an output for each type of pixel on the sensor array comprises measuring a response of that type of pixel.
19. 1. A method for optimizing the alignment of X different types of pixels on an image sensor array for a camera capable of performing N different camera tasks, where X and N are integers greater than 1, the method comprising: simulating pixel responses to training images such that for each training image X sensor images are output, each sensor image uniquely corresponding to one and only one of the X different types located at each possible location on the image sensor array; generating a pixel alignment pattern by distributing the X different types of pixels over the image sensor array locations as a 2D tessellation in a trainable pixel alignment layer; subsampling each of the X sensor images according to the pixel alignment pattern to output X corresponding subsampled sensor images; training the pixel alignment layer and a stack of N neural networks simultaneously, each neural network uniquely corresponding to only one of the N different camera tasks, receiving and processing one or more of the X sub-sampled sensor images, evaluating output values reflecting a user-determined degree of performance of each of the N different camera tasks, and making a discrete and distinct selection of the X different types of pixels so that, taking into account weightings within each camera task, a pixel alignment pattern that produces a relatively high overall output value is adjusted as a candidate for an optimal pixel alignment pattern in which the pixel alignment pattern for the N different camera tasks is optimized according to a balance of priorities within each camera task; A method comprising:
20. 20. The method of claim 19, wherein simultaneously training the pixel alignment layer and N neural networks comprises the parallel use of an inner feedback loop that feeds back to each of the N neural networks a sum of weighted output values derived from the N outputs from the stack, taking into account weightings in each camera task, as an output of a combined performance value, and an outer feedback loop that feeds back the combined performance value to the trainable pixel alignment layer.
21. 21. The method of claim 20, wherein the combined performance value output from the stack comprises a total loss value comprising a combination of performance value outputs from each of the N neural networks, each performance value output being a function of the loss value for the corresponding neural network and at least one user-provided weighting parameter for the camera task corresponding to that neural network.
22. 22. The method of claim 21, wherein the function of a loss value and at least one user-provided weighting parameter for one of the neural networks is a linear function of that loss value.
23. 20. The method of claim 19, wherein at least one of the N tasks is primarily dependent on image quality captured by pixels in the sensor array, and at least another of the N tasks is not primarily dependent on image quality captured by pixels in the sensor array.
24. 1. A system for optimizing the alignment of X different types of pixels on an image sensor array for a camera capable of performing N different camera tasks, where X and N are integers greater than 1, the system comprising: a pixel response simulator; a trainable pixel alignment layer; A subsampler and a stack of N trainable neural networks, each neural network uniquely corresponding to one of the N different camera tasks; Including, the pixel response simulator operates on training images; configured to send, for each training image, X sensor images to the sub-sampler, each sensor image uniquely corresponding to one and only one type of pixel located at each possible location on the image sensor array; the trainable alignment layer generates and sends to a sub-sampler an alignment pattern of pixels where the X different types of pixels are distributed over the locations of the image sensor array as a 2D tessellation; the subsampler is configured to subsample each of the X sensor images according to the pixel alignment pattern and to send one or more subsampled output images to one or more of the neural networks for processing; The N neural networks and the pixel alignment layer are trained simultaneously using an inner feedback loop and an outer feedback loop, respectively, to evaluate output values reflecting the degree of performance of each of the N camera tasks as determined by the user, and perform individual category selection of the X different types of pixels so that, taking into account the weighting in each camera task, a pixel alignment pattern that results in a relatively high overall output value is adjusted as a candidate for an optimal pixel alignment pattern in which the pixel alignment pattern is optimized for performance of the N camera tasks according to a balance of priorities in each camera task.
Citation Information
Patent Citations
Method and related device for processing fundus images based on depth learning
CN109427052A
Configurable pixel array system and method
US20080278610A1
Sparse infrared pixel design for image sensors
US20210297607A1
Apparatus and method for evaluating a quality of image capture of a scene
WO2021048107A1